W06 模型宿主、角色、预算与流式合同:模型宿主合同、角色策略冻结、预算预留结算与流式回合证据。
按 R2 串行阶段整理提交;包内文件为该阶段交付(含后续小增量),状态以工作包清单为准。
This commit is contained in:
parent
e99a952dd0
commit
367d01db21
17
.agent/角色/写手.md
Normal file
17
.agent/角色/写手.md
Normal file
@ -0,0 +1,17 @@
|
||||
---
|
||||
id: writer
|
||||
name: 写手
|
||||
category: role
|
||||
contract_version: 1
|
||||
description: 依据确认规划和固定来源生成正文候选,区分探索与生成阶段。
|
||||
---
|
||||
|
||||
你是写手,负责把本次允许的创作材料写成正文候选。
|
||||
|
||||
先核对任务阶段、写作目标、确认规划、来源范围和保护要求。材料缺失或彼此矛盾时说明具体缺口,不擅自填补正式事实。
|
||||
|
||||
探索阶段只使用任务开放的只读工具,定位人物状态、因果、场景和衔接证据,按输入合同提交探索结果,不提前生成正文。生成阶段只使用已经冻结的材料,不调用工具,一次给出本次要求的完整候选。
|
||||
|
||||
遵守作者选择的声音、方法和篇幅要求;用具体行为、场景细节与有意图的对白推进叙事。写法由本次方法与声音约束决定,不把单一技法强加给所有作品。
|
||||
|
||||
输出结构以任务提供的固定合同为准,不自造身份、来源版本、哈希、检查结论或采纳决定。候选不自动成为正式正文,也不自动确认其中的事实。
|
||||
15
.agent/角色/检测员.md
Normal file
15
.agent/角色/检测员.md
Normal file
@ -0,0 +1,15 @@
|
||||
---
|
||||
id: detector
|
||||
name: 检测员
|
||||
category: role
|
||||
contract_version: 1
|
||||
description: 对固定候选和证据给出可定位的问题、覆盖范围与不确定项。
|
||||
---
|
||||
|
||||
你是检测员,负责核对本次候选与所给依据是否一致。
|
||||
|
||||
先固定候选版本、检查范围和适用规则。逐项核对规划约束、历史事实、能力条件、角色知情范围与作者要求;工具只回读允许的证据,不扩大来源范围。
|
||||
|
||||
每个问题给出候选中的具体引文、位置和支持判断的依据。区分明确冲突、证据不足与本项未见问题,说明未覆盖的部分。不能将缺少证据等同通过。
|
||||
|
||||
按任务输出合同交付报告,不擅自修改候选、不推进叙事状态,不替作者采纳。报告和问题只对所检查的版本有效。
|
||||
7
.agent/角色/目录.md
Normal file
7
.agent/角色/目录.md
Normal file
@ -0,0 +1,7 @@
|
||||
| 名称 | 相对地址 | 内容描述 | 使用场景 | 使用要求 |
|
||||
|------|----------|----------|----------|----------|
|
||||
| 写手 | 写手.md | 两阶段写作职责及候选边界 | 新版写作任务装配 | 使用固定版本和任务授予的输入,不自行外发 |
|
||||
| 规划员 | 规划员.md | 规划因果、层级及作者决定边界 | 新版规划任务装配 | 结构由任务提供,输出为候选 |
|
||||
| 知识抽取员 | 知识抽取员.md | 来源、身份和增量提案职责 | 新版分析与事实抽取 | 使用有效类型和固定来源,不直接确认事实 |
|
||||
| 检测员 | 检测员.md | 问题、证据和覆盖边界 | 新版一致性与质量检查 | 结论绑定被检版本,不修改候选 |
|
||||
| 评委 | 评委.md | 独立评价和匿名判断职责 | 新版评分与比较 | 不读取隔离身份与答案,不代替作者批准 |
|
||||
15
.agent/角色/知识抽取员.md
Normal file
15
.agent/角色/知识抽取员.md
Normal file
@ -0,0 +1,15 @@
|
||||
---
|
||||
id: extractor
|
||||
name: 知识抽取员
|
||||
category: role
|
||||
contract_version: 1
|
||||
description: 从固定来源提取有证据的分析或事实提案,保留时间、身份与不确定性。
|
||||
---
|
||||
|
||||
你是知识抽取员,负责从本次固定来源中提取可核查的内容。
|
||||
|
||||
先区分参考作品分析与作者正文的事实增量。依据输入中的有效类型、字段结构、来源范围和确认状态工作,不自行限定类型总数或把参考作品内容写入作者作品。
|
||||
|
||||
为每项结论保留具体文本依据和位置;区分当时已发生的事实、角色所知、叙述推断及未证实信息。回读已有对象时只使用任务许可的接口和范围,身份疑点不能靠相似名字直接合并。
|
||||
|
||||
缺证据就保留未知或不确定,不补写来源没有表达的细节。变化作为有来源的增量提案,不覆盖历史记录。按固定输出合同交付,不宣称提案已获作者确认,不直接写入正式事实。
|
||||
15
.agent/角色/规划员.md
Normal file
15
.agent/角色/规划员.md
Normal file
@ -0,0 +1,15 @@
|
||||
---
|
||||
id: planner
|
||||
name: 规划员
|
||||
category: role
|
||||
contract_version: 1
|
||||
description: 将作者意图和允许事实组织成有因果、可执行的规划候选。
|
||||
---
|
||||
|
||||
你是规划员,负责明确故事中要发生什么、为什么发生,以及仍需作者选择的事项。
|
||||
|
||||
先核对规划层级、已确认的上层决定、可用事实及本次允许改动的范围。下层规划应承接上层约束;遇到冲突先呈现取舍,不能默改作者已经确认的路线。
|
||||
|
||||
交代触发条件、参与者、行动、结果方向和必要的伏笔安排。保留未知项与假设,区分作者决定、来源事实和规划建议。探索工具只读取任务许可的证据,不读取被排除的未来内容或私人稿。
|
||||
|
||||
按本次固定结构输出规划候选,不把具体字段名或篇幅比例固化为所有作品共用规则。规划不替正文,不替作者批准,也不直接改写正式内容。
|
||||
15
.agent/角色/评委.md
Normal file
15
.agent/角色/评委.md
Normal file
@ -0,0 +1,15 @@
|
||||
---
|
||||
id: judge
|
||||
name: 评委
|
||||
category: role
|
||||
contract_version: 1
|
||||
description: 依据固定评价条件独立判断候选质量,以引文支撑逐维结论。
|
||||
---
|
||||
|
||||
你是评委,负责给出独立、可复核的质量判断。
|
||||
|
||||
先核对评价对象、范围、量表与比较条件。逐维判断情节、人物、语言和场景等本次要求的项目,用具体引文解释结论,不以篇幅、术语数量或笼统好恶代替判断。
|
||||
|
||||
匿名评价只能使用固定的匿名输入,不打听或推测模型、供应商、实验臂、作者身份或标准答案;没有工具许可就不回读外部资料。证据不足或条件不一致时明确指出,不能虚构确定分数。
|
||||
|
||||
按本次固定结构交付评分、比较、依据及分歧。评价不代替确定性门禁、作者决定或正式提交;保留文学判断与自动化结果的区别。
|
||||
105
src/muse/任务运行/工具调用.py
Normal file
105
src/muse/任务运行/工具调用.py
Normal file
@ -0,0 +1,105 @@
|
||||
"""只读工具的冻结范围与回执;查询实现由资料所属模块登记。"""
|
||||
|
||||
from collections.abc import Callable
|
||||
from dataclasses import dataclass
|
||||
from typing import Any
|
||||
|
||||
from jsonschema import Draft202012Validator
|
||||
from jsonschema.exceptions import ValidationError
|
||||
|
||||
from muse.任务运行.执行合同 import 工具请求, 模型协议错误
|
||||
from muse.任务运行.模型 import 内容哈希
|
||||
from muse.任务运行.模型调用 import 校验输出合同
|
||||
|
||||
|
||||
@dataclass(frozen=True, slots=True)
|
||||
class 工具范围:
|
||||
任务ID: str
|
||||
作品ID: str | None
|
||||
来源ID: tuple[str, ...]
|
||||
截止位置: int | None
|
||||
运行用途: str
|
||||
内容用途: str
|
||||
|
||||
|
||||
@dataclass(frozen=True, slots=True)
|
||||
class 工具来源:
|
||||
来源ID: str
|
||||
数据版本: str
|
||||
结构哈希: str
|
||||
投影版本: str
|
||||
|
||||
|
||||
@dataclass(frozen=True, slots=True)
|
||||
class 工具结果:
|
||||
内容: Any
|
||||
来源: tuple[工具来源, ...]
|
||||
剔除原因: tuple[str, ...] = ()
|
||||
|
||||
|
||||
@dataclass(frozen=True, slots=True)
|
||||
class 工具定义:
|
||||
名称: str
|
||||
说明: str
|
||||
参数合同: dict
|
||||
读取: Callable[[工具范围, dict], 工具结果]
|
||||
|
||||
|
||||
class 只读工具集:
|
||||
def __init__(
|
||||
self,
|
||||
登记: tuple[工具定义, ...],
|
||||
允许: tuple[str, ...],
|
||||
范围: 工具范围,
|
||||
记录: Callable[[dict], None],
|
||||
) -> None:
|
||||
索引 = {项.名称: 项 for 项 in 登记}
|
||||
if len(索引) != len(登记) or set(允许) - 索引.keys():
|
||||
raise 模型协议错误("工具未登记、重复或所需工具加载失败")
|
||||
self.工具 = {名: 索引[名] for 名 in 允许}
|
||||
self.范围 = 范围
|
||||
self.记录 = 记录
|
||||
for 定义 in self.工具.values():
|
||||
校验输出合同(定义.参数合同)
|
||||
|
||||
def 调用(self, 请求: 工具请求) -> 工具结果:
|
||||
定义 = self.工具.get(请求.名称)
|
||||
if 定义 is None:
|
||||
raise 模型协议错误("工具不在冻结允许集中")
|
||||
if {
|
||||
"work_id",
|
||||
"as_of",
|
||||
"run_purpose",
|
||||
"content_purpose",
|
||||
"source_scope",
|
||||
"author_id",
|
||||
"decided_by",
|
||||
} & 请求.参数.keys():
|
||||
raise 模型协议错误("模型不得通过参数改变任务身份或范围")
|
||||
try:
|
||||
Draft202012Validator(定义.参数合同).validate(请求.参数)
|
||||
except ValidationError:
|
||||
raise 模型协议错误("工具参数不符合固定查询合同") from None
|
||||
结果 = 定义.读取(self.范围, 请求.参数)
|
||||
if {源.来源ID for 源 in 结果.来源} - set(self.范围.来源ID):
|
||||
raise 模型协议错误("工具返回了冻结范围以外的来源")
|
||||
if any(not all((源.数据版本, 源.结构哈希, 源.投影版本)) for 源 in 结果.来源):
|
||||
raise 模型协议错误("工具来源缺少数据、结构或投影版本")
|
||||
self.记录(
|
||||
{
|
||||
"task_id": self.范围.任务ID,
|
||||
"call_id": 请求.调用ID,
|
||||
"tool": 请求.名称,
|
||||
"arguments_hash": 内容哈希(请求.参数),
|
||||
"source_refs": [
|
||||
{
|
||||
"source_id": 源.来源ID,
|
||||
"revision": 源.数据版本,
|
||||
"schema_hash": 源.结构哈希,
|
||||
"projection_version": 源.投影版本,
|
||||
}
|
||||
for 源 in 结果.来源
|
||||
],
|
||||
}
|
||||
)
|
||||
return 结果
|
||||
115
src/muse/任务运行/执行合同.py
Normal file
115
src/muse/任务运行/执行合同.py
Normal file
@ -0,0 +1,115 @@
|
||||
"""宿主只提供客观执行结果;结构、预算、证据及正式内容由所属模块处理。"""
|
||||
|
||||
from collections.abc import Awaitable, Callable
|
||||
from dataclasses import dataclass
|
||||
from typing import Any, Literal, Protocol
|
||||
|
||||
from muse.共享.错误 import Muse错误
|
||||
|
||||
|
||||
class 模型协议错误(Muse错误):
|
||||
错误码 = "MODEL_PROTOCOL_INVALID"
|
||||
|
||||
|
||||
@dataclass(frozen=True, slots=True)
|
||||
class 模型用量:
|
||||
输入token: int
|
||||
输出token: int
|
||||
缓存读取token: int = 0
|
||||
缓存写入token: int = 0
|
||||
|
||||
def __post_init__(self):
|
||||
if any(
|
||||
type(n) is not int or n < 0
|
||||
for n in (
|
||||
self.输入token,
|
||||
self.输出token,
|
||||
self.缓存读取token,
|
||||
self.缓存写入token,
|
||||
)
|
||||
):
|
||||
raise 模型协议错误("模型用量必须是非负整数;缺少用量应记未知")
|
||||
|
||||
|
||||
@dataclass(frozen=True, slots=True)
|
||||
class 工具请求:
|
||||
调用ID: str
|
||||
名称: str
|
||||
参数: dict[str, Any]
|
||||
|
||||
|
||||
@dataclass(frozen=True, slots=True)
|
||||
class 模型结果:
|
||||
状态: Literal["completed", "incomplete", "failed", "cancelled", "tool_calls"]
|
||||
文本: str
|
||||
实际模型: str | None
|
||||
用量: 模型用量 | None
|
||||
响应ID: str | None = None
|
||||
工具调用: tuple[工具请求, ...] = ()
|
||||
失败码: str | None = None
|
||||
回放协议: str | None = None
|
||||
回放内容: tuple[dict[str, Any], ...] = ()
|
||||
|
||||
@classmethod
|
||||
def 从记录(cls, 数据: dict) -> "模型结果":
|
||||
return cls(
|
||||
**{
|
||||
**数据,
|
||||
"用量": 模型用量(**数据["用量"]) if 数据["用量"] else None,
|
||||
"工具调用": tuple(工具请求(**t) for t in 数据["工具调用"]),
|
||||
"回放内容": tuple(数据["回放内容"]),
|
||||
},
|
||||
)
|
||||
|
||||
|
||||
@dataclass(frozen=True, slots=True)
|
||||
class 工具回传:
|
||||
请求: 工具请求
|
||||
内容: str
|
||||
|
||||
|
||||
@dataclass(frozen=True, slots=True)
|
||||
class 回合记录:
|
||||
模型: 模型结果
|
||||
工具: tuple[工具回传, ...] = ()
|
||||
|
||||
|
||||
@dataclass(frozen=True, slots=True)
|
||||
class 模型请求:
|
||||
调用ID: str
|
||||
provider: str
|
||||
model: str
|
||||
系统提示: str
|
||||
用户输入: str
|
||||
输出合同: dict[str, Any]
|
||||
最大输出token: int
|
||||
总期限秒: float
|
||||
thinking: str | None = None
|
||||
工具: tuple[dict[str, Any], ...] = ()
|
||||
允许实际模型: tuple[str, ...] = ()
|
||||
历史: tuple[回合记录, ...] = ()
|
||||
|
||||
def __post_init__(self):
|
||||
if not all((self.调用ID, self.provider, self.model, self.系统提示, self.用户输入)):
|
||||
raise 模型协议错误("调用身份、显式模型与输入不能为空")
|
||||
if self.最大输出token < 1 or self.总期限秒 <= 0:
|
||||
raise 模型协议错误("调用必须具有正数输出上限和总期限")
|
||||
|
||||
|
||||
class 已准备模型调用(Protocol):
|
||||
async def 调用(self, 收到片段: Callable[[str], None] | None = None) -> 模型结果: ...
|
||||
|
||||
|
||||
class 模型宿主(Protocol):
|
||||
def 准备(self, 请求: 模型请求) -> 已准备模型调用:
|
||||
"""核对本地配置并固定请求;准备过程不得访问模型或预留预算。"""
|
||||
...
|
||||
|
||||
|
||||
class 角色循环(Protocol):
|
||||
async def 运行(
|
||||
self,
|
||||
请求: 模型请求,
|
||||
模型回合: Callable[[], Awaitable[模型结果]],
|
||||
工具回合: Callable[[工具请求], Awaitable[str]],
|
||||
) -> 模型结果: ...
|
||||
354
src/muse/任务运行/模型调用.py
Normal file
354
src/muse/任务运行/模型调用.py
Normal file
@ -0,0 +1,354 @@
|
||||
"""模型调用共用角色、预算、证据和明确终态;业务 owner 决定候选的去留。"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import asyncio
|
||||
import hashlib
|
||||
import json
|
||||
from dataclasses import asdict, dataclass, replace
|
||||
from decimal import Decimal
|
||||
from typing import TYPE_CHECKING, Any, Protocol
|
||||
|
||||
from jsonschema import Draft202012Validator
|
||||
from jsonschema.exceptions import SchemaError, ValidationError
|
||||
|
||||
from muse.任务运行.任务领取 import 校验领取
|
||||
from muse.任务运行.原文生命周期 import 原文服务
|
||||
from muse.任务运行.存储 import 任务存储
|
||||
from muse.任务运行.执行合同 import 模型协议错误, 模型宿主, 模型结果, 模型请求
|
||||
from muse.任务运行.模型 import 执行上下文, 状态冲突
|
||||
from muse.任务运行.角色策略 import 角色策略目录
|
||||
from muse.任务运行.运行证据 import 证据服务
|
||||
from muse.任务运行.配置版本 import 配置快照
|
||||
from muse.任务运行.预算管理 import 预算不足, 预算管理
|
||||
|
||||
if TYPE_CHECKING:
|
||||
from muse.任务运行.接口 import 任务服务
|
||||
|
||||
|
||||
def 校验输出合同(合同: dict) -> None:
|
||||
def 读取(值):
|
||||
if isinstance(值, dict):
|
||||
for 键, 内容 in 值.items():
|
||||
if 键 in {"$ref", "$dynamicRef"} and not 内容.startswith("#"):
|
||||
raise 模型协议错误("输出合同只能引用当前固定结构,不能联网读取定义")
|
||||
读取(内容)
|
||||
elif isinstance(值, list):
|
||||
for 项 in 值:
|
||||
读取(项)
|
||||
|
||||
读取(合同)
|
||||
try:
|
||||
Draft202012Validator.check_schema(合同)
|
||||
except SchemaError:
|
||||
raise 模型协议错误("输出结构合同无效") from None
|
||||
|
||||
|
||||
def 校验模型输出(请求: 模型请求, 结果: 模型结果) -> Any:
|
||||
校验模型回合(请求, 结果)
|
||||
if 结果.状态 != "completed":
|
||||
raise 模型协议错误("模型没有给出明确完成的最终输出", 上下文={"state": 结果.状态})
|
||||
if 结果.实际模型 not in (请求.允许实际模型 or (请求.model,)):
|
||||
raise 模型协议错误("实际模型与冻结模型策略不符")
|
||||
校验输出合同(请求.输出合同)
|
||||
try:
|
||||
值 = json.loads(结果.文本)
|
||||
Draft202012Validator(请求.输出合同).validate(值)
|
||||
except (ValueError, ValidationError):
|
||||
raise 模型协议错误("模型输出不符合固定结构合同") from None
|
||||
return 值
|
||||
|
||||
|
||||
def 校验模型回合(请求: 模型请求, 结果: 模型结果) -> None:
|
||||
if 结果.状态 not in {"completed", "tool_calls"}:
|
||||
raise 模型协议错误("模型回合缺少明确终态", 上下文={"state": 结果.状态})
|
||||
if 结果.实际模型 not in (请求.允许实际模型 or (请求.model,)):
|
||||
raise 模型协议错误("实际模型与冻结模型策略不符")
|
||||
if 结果.状态 == "tool_calls":
|
||||
if not 结果.工具调用 or {t.名称 for t in 结果.工具调用} - {t["name"] for t in 请求.工具}:
|
||||
raise 模型协议错误("模型回合请求了未登记工具")
|
||||
elif 结果.工具调用:
|
||||
raise 模型协议错误("模型终态与工具请求不一致")
|
||||
|
||||
|
||||
def 请求字节(请求: 模型请求) -> bytes:
|
||||
参数 = asdict(请求)
|
||||
# 后验模型别名来自任务冻结策略,不是发给模型的请求参数。
|
||||
参数.pop("允许实际模型")
|
||||
return json.dumps(参数, ensure_ascii=False, sort_keys=True, separators=(",", ":")).encode()
|
||||
|
||||
|
||||
class 调用计价(Protocol):
|
||||
"""装配注入已登记版本的费用规则;无法可靠计价时返回未知,不能推定免费。"""
|
||||
|
||||
版本: str
|
||||
|
||||
def 金额(self, 结果: 模型结果) -> Decimal | None: ...
|
||||
|
||||
|
||||
@dataclass(frozen=True)
|
||||
class 模型交付:
|
||||
内容: Any
|
||||
证据回执: dict
|
||||
调用ID: str
|
||||
|
||||
|
||||
@dataclass(frozen=True)
|
||||
class 回合交付:
|
||||
结果: 模型结果
|
||||
证据回执: dict
|
||||
调用ID: str
|
||||
|
||||
|
||||
class 模型执行器:
|
||||
def __init__(
|
||||
self,
|
||||
运行: 任务服务,
|
||||
预算: 预算管理,
|
||||
原文: 原文服务,
|
||||
证据: 证据服务,
|
||||
角色: 角色策略目录,
|
||||
宿主: 模型宿主,
|
||||
计价: 调用计价,
|
||||
*,
|
||||
固定配置: 配置快照 | None = None,
|
||||
) -> None:
|
||||
self.运行, self.预算, self.原文, self.证据 = 运行, 预算, 原文, 证据
|
||||
self.角色, self.宿主, self.计价 = 角色, 宿主, 计价
|
||||
self.固定配置 = 固定配置
|
||||
|
||||
async def 执行(
|
||||
self,
|
||||
上下文: 执行上下文,
|
||||
请求: 模型请求,
|
||||
*,
|
||||
阶段: str,
|
||||
原文授权ID: str,
|
||||
临时租约ID: str | None = None,
|
||||
) -> 模型交付:
|
||||
"""保留方式由明确批准决定;临时字节不进入普通PG运行证据。"""
|
||||
回合 = await self._执行回合(
|
||||
上下文,
|
||||
请求,
|
||||
阶段=阶段,
|
||||
原文授权ID=原文授权ID,
|
||||
校验最终=True,
|
||||
临时租约ID=临时租约ID,
|
||||
)
|
||||
# 回合中已校验并保存;这里只组装单回合角色的业务交付。
|
||||
return 模型交付(json.loads(回合.结果.文本), 回合.证据回执, 回合.调用ID)
|
||||
|
||||
async def 执行回合(
|
||||
self, 上下文: 执行上下文, 请求: 模型请求, *, 阶段: str, 原文授权ID: str
|
||||
) -> 回合交付:
|
||||
return await self._执行回合(上下文, 请求, 阶段=阶段, 原文授权ID=原文授权ID, 校验最终=False)
|
||||
|
||||
def 核对执行合同(self, 上下文: 执行上下文, 请求: 模型请求, 阶段: str):
|
||||
领取 = 上下文.领取
|
||||
当前任务 = self.运行.读取任务(领取.任务ID)
|
||||
步骤 = next(s for s in 当前任务.流程.步骤 if s.步骤ID == 领取.步骤ID)
|
||||
if 步骤.角色 is None:
|
||||
raise 模型协议错误("确定性步骤不能临时调用模型")
|
||||
if (
|
||||
not self.角色.资源发布身份
|
||||
or 当前任务.冻结输入["资源发布身份"] != self.角色.资源发布身份
|
||||
):
|
||||
raise 模型协议错误("任务冻结资源与当前执行构建不一致,不能自动使用新版资源")
|
||||
策略 = self.角色.冻结(
|
||||
步骤.角色,
|
||||
provider=请求.provider,
|
||||
model=请求.model,
|
||||
thinking=请求.thinking,
|
||||
工具=tuple(t["name"] for t in 请求.工具),
|
||||
阶段=阶段,
|
||||
)
|
||||
if (
|
||||
当前任务.冻结输入["角色策略版本"] != 策略.策略版本
|
||||
or tuple(t["name"] for t in 请求.工具) != 步骤.工具
|
||||
):
|
||||
raise 模型协议错误("调用角色版本或工具不符合任务冻结合同")
|
||||
if self.固定配置 is not None:
|
||||
设定 = self.固定配置.内容.角色配置.get(策略.角色)
|
||||
if 设定 is None or (
|
||||
请求.provider,
|
||||
请求.model,
|
||||
请求.thinking,
|
||||
tuple(t["name"] for t in 请求.工具),
|
||||
) != (设定["provider"], 设定["model"], 设定["thinking"], tuple(设定.get("tools", []))):
|
||||
raise 模型协议错误("角色调用与任务固定的运行配置不一致")
|
||||
if self.计价.版本 != self.固定配置.内容.计价版本:
|
||||
raise 模型协议错误("费用规则与任务固定配置版本不一致")
|
||||
if (
|
||||
设定.get("input_schema", 步骤.输入合同) != 步骤.输入合同
|
||||
or 设定.get("output_schema", 步骤.输出合同) != 步骤.输出合同
|
||||
):
|
||||
raise 模型协议错误("运行配置声明的输入输出合同与固定步骤不一致")
|
||||
请求 = replace(请求, 允许实际模型=策略.允许实际模型)
|
||||
校验输出合同(请求.输出合同)
|
||||
self.运行.重验执行范围(领取)
|
||||
return 请求, 当前任务, 策略
|
||||
|
||||
async def _执行回合(
|
||||
self,
|
||||
上下文: 执行上下文,
|
||||
请求: 模型请求,
|
||||
*,
|
||||
阶段: str,
|
||||
原文授权ID: str,
|
||||
校验最终: bool,
|
||||
临时租约ID: str | None = None,
|
||||
) -> 回合交付:
|
||||
请求, 当前任务, 策略 = self.核对执行合同(上下文, 请求, 阶段)
|
||||
领取 = 上下文.领取
|
||||
输入 = 请求字节(请求)
|
||||
输入哈希 = hashlib.sha256(输入).hexdigest()
|
||||
self.原文.检查调用授权(
|
||||
领取.任务ID,
|
||||
原文授权ID,
|
||||
请求.调用ID,
|
||||
输入哈希,
|
||||
"temporary" if 临时租约ID is not None else "persistent",
|
||||
请求.总期限秒 + 30,
|
||||
)
|
||||
if 临时租约ID is not None:
|
||||
self.原文.检查执行余量(领取.任务ID, 临时租约ID, 请求.总期限秒 + 30, 授权ID=原文授权ID)
|
||||
已准备 = self.宿主.准备(请求)
|
||||
元信息 = {
|
||||
"call_id": 请求.调用ID,
|
||||
"model": 请求.model,
|
||||
"provider": 请求.provider,
|
||||
"policy_version": 策略.策略版本,
|
||||
"resource_release_id": 当前任务.冻结输入["资源发布身份"],
|
||||
"pricing_version": self.计价.版本,
|
||||
}
|
||||
if self.固定配置:
|
||||
元信息["runtime_config_hash"] = self.固定配置.内容哈希
|
||||
if 临时租约ID is not None:
|
||||
元信息["raw_lease_id"] = 临时租约ID
|
||||
self.证据.保存(
|
||||
领取.任务ID,
|
||||
领取.尝试ID,
|
||||
"model_input",
|
||||
输入哈希,
|
||||
"completed",
|
||||
元信息,
|
||||
数据=输入 if 临时租约ID is None else None,
|
||||
授权ID=原文授权ID,
|
||||
)
|
||||
if 临时租约ID is not None:
|
||||
self._暂存模型字节(上下文, 请求.调用ID, "model_input", 临时租约ID, 输入哈希, 输入)
|
||||
self.预算.预留(领取, 请求.调用ID, 策略.角色)
|
||||
try:
|
||||
发送 = self.预算.标记已发送(
|
||||
领取,
|
||||
请求.调用ID,
|
||||
最长秒=请求.总期限秒,
|
||||
登记尝试=True,
|
||||
)
|
||||
except Exception:
|
||||
self.预算.取消预留(请求.调用ID)
|
||||
raise
|
||||
if not 发送.允许外发:
|
||||
raise 状态冲突("调用已发出,读取证据或对账,不能重复外发")
|
||||
结果 = None
|
||||
try:
|
||||
async with asyncio.timeout(请求.总期限秒):
|
||||
结果 = await 已准备.调用()
|
||||
finally:
|
||||
# 结构无效、取消和证据保存失败都不抹掉已发生的费用。
|
||||
try:
|
||||
成本 = None if 结果 is None else self.计价.金额(结果)
|
||||
except Exception:
|
||||
成本 = None
|
||||
结算 = self.预算.结算(
|
||||
请求.调用ID,
|
||||
成本,
|
||||
回执ID=结果.响应ID if 结果 and 结果.响应ID else 请求.调用ID,
|
||||
)
|
||||
if 结果 is None:
|
||||
raise 模型协议错误("模型宿主没有返回执行结果")
|
||||
# 工具身份、参数与协议回放块也是响应的一部分,不能只保存可见文本。
|
||||
响应 = json.dumps(
|
||||
asdict(结果), ensure_ascii=False, sort_keys=True, separators=(",", ":")
|
||||
).encode()
|
||||
响应哈希 = hashlib.sha256(响应).hexdigest()
|
||||
self.原文.绑定调用响应(领取.任务ID, 原文授权ID, 请求.调用ID, 输入哈希, 响应哈希)
|
||||
错误 = None
|
||||
try:
|
||||
if 校验最终:
|
||||
校验模型输出(请求, 结果)
|
||||
else:
|
||||
校验模型回合(请求, 结果)
|
||||
except 模型协议错误 as exc:
|
||||
错误 = exc
|
||||
回执 = self.证据.保存(
|
||||
领取.任务ID,
|
||||
领取.尝试ID,
|
||||
"model_response",
|
||||
响应哈希,
|
||||
"completed" if 错误 is None else "failed",
|
||||
{
|
||||
**元信息,
|
||||
"model": 结果.实际模型,
|
||||
"input_tokens": 结果.用量.输入token if 结果.用量 else None,
|
||||
"output_tokens": 结果.用量.输出token if 结果.用量 else None,
|
||||
"error_code": 错误.错误码 if 错误 else None,
|
||||
},
|
||||
数据=响应 if 临时租约ID is None else None,
|
||||
授权ID=原文授权ID,
|
||||
)
|
||||
if 临时租约ID is not None:
|
||||
self._暂存模型字节(上下文, 请求.调用ID, "model_response", 临时租约ID, 响应哈希, 响应)
|
||||
if 结算.状态 == "unknown":
|
||||
raise 预算不足("响应证据已保存,但成本未知,必须对账后再继续")
|
||||
self._确认保存(上下文, 请求.调用ID, 回执["evidence_id"], 临时租约ID=临时租约ID)
|
||||
if 错误:
|
||||
raise 错误
|
||||
if 结算.超预算:
|
||||
raise 预算不足("实际成本已登记并超过批准预算")
|
||||
return 回合交付(结果, 回执, 请求.调用ID)
|
||||
|
||||
def _暂存模型字节(
|
||||
self, 上下文: 执行上下文, 调用ID: str, 种类: str, 租约ID: str, 哈希: str, 数据: bytes
|
||||
) -> None:
|
||||
try:
|
||||
self.原文.写入(上下文.领取.任务ID, 租约ID, 哈希, 数据)
|
||||
except Exception:
|
||||
self.证据.保存(
|
||||
上下文.领取.任务ID,
|
||||
上下文.领取.尝试ID,
|
||||
"failure",
|
||||
哈希,
|
||||
"failed",
|
||||
{
|
||||
"call_id": f"{调用ID}:raw:{种类}",
|
||||
"error_code": "RAW_TEMP_WRITE_FAILED",
|
||||
"raw_lease_id": 租约ID,
|
||||
},
|
||||
)
|
||||
raise
|
||||
|
||||
def _确认保存(
|
||||
self, 上下文: 执行上下文, 调用ID: str, 证据ID: str, *, 临时租约ID: str | None = None
|
||||
) -> None:
|
||||
if 临时租约ID is not None:
|
||||
回执 = self.证据.读取回执(上下文.领取.任务ID, 证据ID)
|
||||
self.原文.读取(上下文.领取.任务ID, 临时租约ID, 回执["content_hash"])
|
||||
with self.运行.数据库.连接() as 连, 连.transaction():
|
||||
存储 = 任务存储(连, self.运行.数据库.用途)
|
||||
_, 尝试 = 校验领取(存储, 上下文.领取)
|
||||
证据 = 存储.查询(
|
||||
"SELECT content,metadata FROM {s}.muse_runtime_evidence WHERE task_id=%s "
|
||||
"AND attempt_id=%s AND evidence_id=%s",
|
||||
(上下文.领取.任务ID, 上下文.领取.尝试ID, 证据ID),
|
||||
).fetchone()
|
||||
if 尝试["call_reference"] != 调用ID or 证据 is None:
|
||||
raise 状态冲突("调用或完整证据没有对应当前尝试")
|
||||
if 证据["content"] is None and (
|
||||
临时租约ID is None or 证据["metadata"].get("raw_lease_id") != 临时租约ID
|
||||
):
|
||||
raise 状态冲突("完整响应未持久保存,也未关联已核对临时租约")
|
||||
存储.查询(
|
||||
"UPDATE {s}.muse_attempt SET call_state='saved' WHERE attempt_id=%s",
|
||||
(上下文.领取.尝试ID,),
|
||||
)
|
||||
229
src/muse/任务运行/角色会话.py
Normal file
229
src/muse/任务运行/角色会话.py
Normal file
@ -0,0 +1,229 @@
|
||||
"""S02拥有角色对话、逐回合证据与工具来源;宿主只驱动受限循环。"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import hashlib
|
||||
import json
|
||||
import uuid
|
||||
from dataclasses import asdict, replace
|
||||
from typing import TYPE_CHECKING
|
||||
|
||||
from muse.任务运行.工具调用 import 只读工具集, 工具定义, 工具范围
|
||||
from muse.任务运行.执行合同 import 回合记录, 工具回传, 工具请求, 模型协议错误, 模型结果, 模型请求
|
||||
from muse.任务运行.模型 import 事件类型, 执行上下文
|
||||
from muse.任务运行.模型调用 import 回合交付, 校验模型输出, 模型交付, 模型执行器, 请求字节
|
||||
from muse.任务运行.角色策略 import 角色策略目录
|
||||
|
||||
if TYPE_CHECKING:
|
||||
from muse.任务运行.执行合同 import 角色循环
|
||||
|
||||
|
||||
def 组装角色请求(角色: 角色策略目录, 身份: str, 请求: 模型请求) -> 模型请求:
|
||||
"""在作者批准请求哈希之前固定职责和动态输出合同。"""
|
||||
return replace(
|
||||
请求,
|
||||
系统提示=角色.读取角色提示(身份)
|
||||
+ "\n\n本次任务:\n"
|
||||
+ 请求.系统提示
|
||||
+ "\n\n最终输出结构:\n"
|
||||
+ json.dumps(请求.输出合同, ensure_ascii=False, sort_keys=True),
|
||||
)
|
||||
|
||||
|
||||
class 角色会话:
|
||||
def __init__(
|
||||
self,
|
||||
执行器: 模型执行器,
|
||||
上下文: 执行上下文,
|
||||
请求: 模型请求,
|
||||
*,
|
||||
阶段: str,
|
||||
会话授权ID: str,
|
||||
工具登记: tuple[工具定义, ...],
|
||||
):
|
||||
请求, _, _ = 执行器.核对执行合同(上下文, 请求, 阶段)
|
||||
执行器.运行.重验执行范围(上下文.领取)
|
||||
self.执行器, self.上下文, self.请求 = 执行器, 上下文, 请求
|
||||
self.阶段, self.会话ID = 阶段, 会话授权ID
|
||||
self.初始哈希 = hashlib.sha256(请求字节(请求)).hexdigest()
|
||||
self.批准 = 执行器.原文.核对会话(上下文.领取, 会话授权ID, self.初始哈希, 阶段)
|
||||
任务 = 执行器.运行.读取任务(上下文.领取.任务ID)
|
||||
步骤 = next(s for s in 任务.流程.步骤 if s.步骤ID == 上下文.领取.步骤ID)
|
||||
if (
|
||||
请求.历史
|
||||
or not 步骤.角色
|
||||
or not 请求.系统提示.startswith(
|
||||
执行器.角色.读取角色提示(步骤.角色) + "\n\n本次任务:\n"
|
||||
)
|
||||
):
|
||||
raise 模型协议错误("角色会话必须从固定职责和批准的初始输入开始")
|
||||
self.角色ID = 步骤.角色
|
||||
固定 = 任务.冻结输入
|
||||
来源 = 固定["冻结上下文"]["source_scope"]
|
||||
self.工具 = 只读工具集(
|
||||
工具登记,
|
||||
步骤.工具,
|
||||
工具范围(
|
||||
任务.任务ID,
|
||||
固定["作品ID"],
|
||||
tuple(来源.get("source_ids", [])),
|
||||
来源.get("as_of"),
|
||||
固定["执行用途"],
|
||||
固定["内容用途"],
|
||||
),
|
||||
self._记录工具读取,
|
||||
)
|
||||
描述 = tuple(
|
||||
{"name": t.名称, "description": t.说明, "parameters": t.参数合同}
|
||||
for t in self.工具.工具.values()
|
||||
)
|
||||
if 描述 != 请求.工具:
|
||||
raise 模型协议错误("请求工具描述与实际安装的只读工具不一致")
|
||||
self.历史: list[回合记录] = []
|
||||
self.证据: list[str] = []
|
||||
self.当前: 回合交付 | None = None
|
||||
self.待工具: dict[str, 工具请求] = {}
|
||||
self.工具回传: list[工具回传] = []
|
||||
self.工具次数 = 0
|
||||
self.结束 = False
|
||||
self.已有证据 = 执行器.原文.读取会话证据(上下文.领取, 会话授权ID, self.初始哈希, 阶段)
|
||||
|
||||
def _记录工具读取(self, 载荷: dict) -> None:
|
||||
self.执行器.运行.追加事件(
|
||||
self.上下文.领取, 事件类型.工具读取, 载荷, 事件ID=str(uuid.uuid4())
|
||||
)
|
||||
|
||||
async def 模型回合(self) -> 模型结果:
|
||||
if self.结束 or self.待工具 or self.当前 and self.当前.结果.状态 != "tool_calls":
|
||||
raise 模型协议错误("当前会话不允许启动下一模型回合")
|
||||
self.执行器.运行.重验执行范围(self.上下文.领取)
|
||||
self.执行器.原文.核对会话(self.上下文.领取, self.会话ID, self.初始哈希, self.阶段)
|
||||
if self.当前:
|
||||
self.历史.append(回合记录(self.当前.结果, tuple(self.工具回传)))
|
||||
self.工具回传 = []
|
||||
次序 = len(self.历史) + 1
|
||||
请求 = (
|
||||
self.请求
|
||||
if 次序 == 1
|
||||
else replace(
|
||||
self.请求, 调用ID=f"{self.请求.调用ID}:round:{次序}", 历史=tuple(self.历史)
|
||||
)
|
||||
)
|
||||
哈希 = hashlib.sha256(请求字节(请求)).hexdigest()
|
||||
旧 = self.已有证据.get(请求.调用ID)
|
||||
if 旧:
|
||||
self.执行器.原文.检查调用授权(
|
||||
self.上下文.领取.任务ID, 旧["authorization_id"], 请求.调用ID, 哈希, "persistent", 0
|
||||
)
|
||||
if 旧["derivation_kind"] != "model" or 旧["call_request_hash"] != 哈希:
|
||||
raise 模型协议错误("已保存回合与当前固定对话不一致")
|
||||
self.当前 = 回合交付(
|
||||
模型结果.从记录(json.loads(旧["content"])), self._旧回执(旧), 请求.调用ID
|
||||
)
|
||||
else:
|
||||
授权 = self.执行器.原文.派生会话保留(
|
||||
self.上下文.领取,
|
||||
self.会话ID,
|
||||
请求.调用ID,
|
||||
哈希,
|
||||
初始请求哈希=self.初始哈希,
|
||||
阶段=self.阶段,
|
||||
种类="model",
|
||||
前序证据=tuple(self.证据),
|
||||
)
|
||||
self.当前 = await self.执行器.执行回合(
|
||||
self.上下文, 请求, 阶段=self.阶段, 原文授权ID=授权
|
||||
)
|
||||
self.证据.append(self.当前.证据回执["evidence_id"])
|
||||
工具 = self.当前.结果.工具调用
|
||||
self.待工具 = {t.调用ID: t for t in 工具}
|
||||
if len(self.待工具) != len(工具):
|
||||
raise 模型协议错误("同一回合的工具调用身份重复")
|
||||
return self.当前.结果
|
||||
|
||||
async def 工具回合(self, 请求: 工具请求) -> str:
|
||||
if self.结束:
|
||||
raise 模型协议错误("角色会话已结束")
|
||||
self.执行器.运行.重验执行范围(self.上下文.领取)
|
||||
self.执行器.原文.核对会话(self.上下文.领取, self.会话ID, self.初始哈希, self.阶段)
|
||||
if self.待工具.get(请求.调用ID) != 请求 or self.工具次数 >= self.批准["max_tool_calls"]:
|
||||
raise 模型协议错误("工具请求未由当前模型提出或超过批准次数")
|
||||
引用 = f"{self.会话ID}:tool:{请求.调用ID}"
|
||||
旧 = self.已有证据.get(引用)
|
||||
if 旧:
|
||||
self.执行器.原文.检查调用授权(
|
||||
self.上下文.领取.任务ID,
|
||||
旧["authorization_id"],
|
||||
引用,
|
||||
旧["call_request_hash"],
|
||||
"persistent",
|
||||
0,
|
||||
)
|
||||
数据 = json.loads(旧["content"])
|
||||
if 旧["derivation_kind"] != "tool" or 数据["request"] != asdict(请求):
|
||||
raise 模型协议错误("已保存工具结果与模型请求不一致")
|
||||
内容 = json.dumps(数据["result"], ensure_ascii=False, sort_keys=True)
|
||||
self._接收工具结果(请求, 内容, 旧["evidence_id"])
|
||||
return 内容
|
||||
结果 = self.工具.调用(请求)
|
||||
完整 = json.dumps(
|
||||
{"request": asdict(请求), "result": asdict(结果)},
|
||||
ensure_ascii=False,
|
||||
sort_keys=True,
|
||||
separators=(",", ":"),
|
||||
).encode()
|
||||
哈希 = hashlib.sha256(完整).hexdigest()
|
||||
授权 = self.执行器.原文.派生会话保留(
|
||||
self.上下文.领取,
|
||||
self.会话ID,
|
||||
引用,
|
||||
哈希,
|
||||
初始请求哈希=self.初始哈希,
|
||||
阶段=self.阶段,
|
||||
种类="tool",
|
||||
前序证据=tuple(self.证据),
|
||||
)
|
||||
回执 = self.执行器.证据.保存(
|
||||
self.上下文.领取.任务ID,
|
||||
self.上下文.领取.尝试ID,
|
||||
"tool_result",
|
||||
哈希,
|
||||
"completed",
|
||||
{"call_id": 引用},
|
||||
数据=完整,
|
||||
授权ID=授权,
|
||||
)
|
||||
内容 = json.dumps(asdict(结果), ensure_ascii=False, sort_keys=True)
|
||||
self._接收工具结果(请求, 内容, 回执["evidence_id"])
|
||||
return 内容
|
||||
|
||||
def _接收工具结果(self, 请求: 工具请求, 内容: str, 证据ID: str) -> None:
|
||||
self.证据.append(证据ID)
|
||||
self.工具回传.append(工具回传(请求, 内容))
|
||||
self.工具次数 += 1
|
||||
del self.待工具[请求.调用ID]
|
||||
|
||||
@staticmethod
|
||||
def _旧回执(记录: dict) -> dict:
|
||||
return {
|
||||
key: 记录[key] for key in ("evidence_id", "content_hash", "outcome", "revision")
|
||||
} | {"retention": "full"}
|
||||
|
||||
async def 执行(self, 宿主: 角色循环) -> 模型交付:
|
||||
try:
|
||||
结果 = await 宿主.运行(self.请求, self.模型回合, self.工具回合)
|
||||
if self.当前 is None or 结果 != self.当前.结果 or self.待工具:
|
||||
raise 模型协议错误("角色宿主没有交回已确认的最终模型结果")
|
||||
self.执行器.原文.核对会话(self.上下文.领取, self.会话ID, self.初始哈希, self.阶段)
|
||||
实际策略 = self.执行器.角色.冻结(
|
||||
self.角色ID,
|
||||
provider=self.请求.provider,
|
||||
model=self.请求.model,
|
||||
thinking=self.请求.thinking,
|
||||
工具=tuple(t["name"] for t in self.请求.工具),
|
||||
阶段=self.阶段,
|
||||
)
|
||||
请求 = replace(self.请求, 允许实际模型=实际策略.允许实际模型)
|
||||
return 模型交付(校验模型输出(请求, 结果), self.当前.证据回执, self.当前.调用ID)
|
||||
finally:
|
||||
self.结束 = True
|
||||
104
src/muse/任务运行/角色策略.py
Normal file
104
src/muse/任务运行/角色策略.py
Normal file
@ -0,0 +1,104 @@
|
||||
"""固定角色能力与模型映射;资源版本在任务创建时冻结,不隐式降级。"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
from dataclasses import dataclass
|
||||
|
||||
import yaml
|
||||
|
||||
from muse.任务运行.执行合同 import 模型协议错误
|
||||
from muse.资源加载 import 加载清单, 打开资源, 读取能力, 读取能力目录
|
||||
|
||||
|
||||
@dataclass(frozen=True, slots=True)
|
||||
class 角色执行策略:
|
||||
角色: str
|
||||
策略版本: str
|
||||
provider: str
|
||||
model: str
|
||||
thinking: str | None
|
||||
允许实际模型: tuple[str, ...]
|
||||
工具: tuple[str, ...]
|
||||
|
||||
|
||||
class 角色策略目录:
|
||||
def __init__(self, 定义: dict, *, 资源发布身份: str = "", 角色资源: dict | None = None) -> None:
|
||||
self.定义 = 定义
|
||||
self.资源发布身份 = 资源发布身份
|
||||
self.角色资源 = 角色资源 or {}
|
||||
if not 定义.get("version") or set(定义.get("roles", {})) != {
|
||||
"writer",
|
||||
"planner",
|
||||
"judge",
|
||||
"detector",
|
||||
"extractor",
|
||||
}:
|
||||
raise 模型协议错误("角色策略必须登记五类角色及明确版本")
|
||||
if any(
|
||||
定义["roles"][角色].get("model_policy") != "fixed"
|
||||
for 角色 in ("writer", "planner", "judge")
|
||||
):
|
||||
raise 模型协议错误("写手、规划员与评委不能改用降级链策略")
|
||||
|
||||
@classmethod
|
||||
def 从发布包(cls) -> 角色策略目录:
|
||||
清单 = 加载清单()
|
||||
角色 = 读取能力目录()["role"]
|
||||
目录 = cls(
|
||||
yaml.safe_load(打开资源("策略/角色.yaml")),
|
||||
资源发布身份=清单["构建身份"],
|
||||
角色资源=角色,
|
||||
)
|
||||
if set(角色) != set(目录.定义["roles"]):
|
||||
raise 模型协议错误("角色策略与已发布职责目录不一致")
|
||||
return 目录
|
||||
|
||||
def 读取角色提示(self, 角色: str) -> str:
|
||||
身份 = self.定义.get("aliases", {}).get(角色, 角色)
|
||||
条目 = self.角色资源.get(身份)
|
||||
if 条目 is None:
|
||||
raise 模型协议错误("角色职责未从固定发布包装配")
|
||||
return 读取能力("role", 身份, 预期哈希=条目["sha256"])["正文"]
|
||||
|
||||
def 冻结(
|
||||
self,
|
||||
角色: str,
|
||||
*,
|
||||
provider: str,
|
||||
model: str,
|
||||
thinking: str | None = None,
|
||||
工具: tuple[str, ...] = (),
|
||||
阶段: str = "执行",
|
||||
) -> 角色执行策略:
|
||||
规范角色 = self.定义.get("aliases", {}).get(角色, 角色)
|
||||
定义 = self.定义["roles"].get(规范角色)
|
||||
if 定义 is None or not provider or not model:
|
||||
raise 模型协议错误("角色和明确提供方、模型必须登记")
|
||||
if 规范角色 == "writer" and 阶段 not in {"探索", "生成"}:
|
||||
raise 模型协议错误("写手必须显式指定探索或生成阶段")
|
||||
允许 = self.定义["models"][定义["model_policy"]]
|
||||
if model not in 允许:
|
||||
raise 模型协议错误("请求模型不符合角色能力策略")
|
||||
if thinking is not None and thinking not in {
|
||||
"off",
|
||||
"minimal",
|
||||
"low",
|
||||
"medium",
|
||||
"high",
|
||||
"xhigh",
|
||||
"max",
|
||||
}:
|
||||
raise 模型协议错误("推理等级未登记")
|
||||
if 工具 and (not 定义["readonly_tools"] or 规范角色 == "writer" and 阶段 == "生成"):
|
||||
raise 模型协议错误("该角色阶段不允许使用工具")
|
||||
if 角色 == "blind_judge" and 工具:
|
||||
raise 模型协议错误("盲评只使用冻结匿名输入")
|
||||
return 角色执行策略(
|
||||
规范角色,
|
||||
self.定义["version"],
|
||||
provider,
|
||||
model,
|
||||
thinking,
|
||||
tuple(self.定义.get("actual_model_ids", {}).get(model, [model])),
|
||||
工具,
|
||||
)
|
||||
389
src/muse/任务运行/配置版本.py
Normal file
389
src/muse/任务运行/配置版本.py
Normal file
@ -0,0 +1,389 @@
|
||||
"""配置版本、验证回执、启用指针与任务冻结副本;不读取未审工作树或凭据值。"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import json
|
||||
import re
|
||||
import uuid
|
||||
from collections.abc import Iterator
|
||||
from contextlib import contextmanager
|
||||
from dataclasses import asdict, dataclass
|
||||
from pathlib import Path
|
||||
from typing import Any, LiteralString, Protocol
|
||||
from urllib.parse import urlsplit
|
||||
|
||||
import psycopg
|
||||
from psycopg import sql
|
||||
from psycopg.rows import dict_row
|
||||
from psycopg.types.json import Jsonb
|
||||
|
||||
from muse.任务运行.存储 import 任务存储
|
||||
from muse.任务运行.模型 import 内容哈希
|
||||
from muse.共享.调用身份 import 用途
|
||||
from muse.共享.错误 import Muse错误
|
||||
from muse.基础设施.数据库.连接 import 数据库工厂
|
||||
|
||||
_角色 = frozenset({"writer", "planner", "extractor", "detector", "judge"})
|
||||
|
||||
|
||||
class 配置版本错误(Muse错误):
|
||||
错误码 = "MUSE_RUNTIME_CONFIG"
|
||||
|
||||
|
||||
@dataclass(frozen=True, slots=True)
|
||||
class 凭据引用:
|
||||
名称: str
|
||||
来源: str
|
||||
位置: str
|
||||
|
||||
def __post_init__(self) -> None:
|
||||
if not self.名称 or self.来源 not in ("环境变量", "受控存储"):
|
||||
raise 配置版本错误("配置凭据只接受明确的引用来源")
|
||||
if self.来源 == "环境变量" and not re.fullmatch(r"[A-Za-z_][A-Za-z0-9_]*", self.位置):
|
||||
raise 配置版本错误("凭据环境变量引用不合法")
|
||||
if self.来源 == "受控存储" and not Path(self.位置).is_absolute():
|
||||
raise 配置版本错误("受控存储引用必须是绝对位置")
|
||||
|
||||
|
||||
@dataclass(frozen=True, slots=True)
|
||||
class 提供方配置:
|
||||
身份: str
|
||||
协议: str
|
||||
地址: str
|
||||
凭据名称: str
|
||||
|
||||
def __post_init__(self) -> None:
|
||||
if not self.身份 or self.协议 not in {"responses", "anthropic", "chat-completions"}:
|
||||
raise 配置版本错误("提供方需要稳定身份和已支持协议")
|
||||
地址 = urlsplit(self.地址)
|
||||
if (
|
||||
地址.scheme not in {"http", "https"}
|
||||
or not 地址.hostname
|
||||
or 地址.username
|
||||
or 地址.password
|
||||
or 地址.fragment
|
||||
or 地址.query
|
||||
):
|
||||
raise 配置版本错误("提供方地址必须为不含认证信息或查询参数的HTTP端点")
|
||||
if not self.凭据名称:
|
||||
raise 配置版本错误("提供方必须绑定具名凭据引用")
|
||||
|
||||
|
||||
@dataclass(frozen=True, slots=True)
|
||||
class 运行配置内容:
|
||||
宿主: str
|
||||
宿主版本: str
|
||||
角色策略版本: str
|
||||
资源发布身份: str
|
||||
预算策略引用: str
|
||||
角色配置: dict[str, dict[str, Any]]
|
||||
凭据: tuple[凭据引用, ...]
|
||||
提供方: tuple[提供方配置, ...]
|
||||
计价版本: str
|
||||
Node路径: str | None = None
|
||||
Pi包目录: str | None = None
|
||||
|
||||
def __post_init__(self) -> None:
|
||||
if self.宿主 not in ("pi", "direct") or not all(
|
||||
(self.宿主版本, self.角色策略版本, self.资源发布身份, self.预算策略引用, self.计价版本)
|
||||
):
|
||||
raise 配置版本错误("配置需要已知宿主以及宿主、角色、资源和预算版本引用")
|
||||
if not self.角色配置 or not set(self.角色配置) <= _角色:
|
||||
raise 配置版本错误("配置只允许显式登记的五类角色")
|
||||
for 策略 in self.角色配置.values():
|
||||
if set(策略) - {
|
||||
"provider",
|
||||
"model",
|
||||
"thinking",
|
||||
"tools",
|
||||
"input_schema",
|
||||
"output_schema",
|
||||
}:
|
||||
raise 配置版本错误("角色配置包含不支持的字段;凭据只能放在引用表")
|
||||
for 字段 in ("provider", "model", "thinking"):
|
||||
if not isinstance(策略.get(字段), str) or not 策略[字段].strip():
|
||||
raise 配置版本错误("角色 provider、model 与 thinking 必须显式配置")
|
||||
for 字段 in ("input_schema", "output_schema"):
|
||||
if 字段 in 策略 and (not isinstance(策略[字段], str) or not 策略[字段]):
|
||||
raise 配置版本错误("结构字段只保存明确的结构版本身份")
|
||||
if "tools" in 策略 and not (
|
||||
isinstance(策略["tools"], list) and all(isinstance(t, str) for t in 策略["tools"])
|
||||
):
|
||||
raise 配置版本错误("角色工具必须是具名身份列表")
|
||||
if len({r.名称 for r in self.凭据}) != len(self.凭据):
|
||||
raise 配置版本错误("凭据引用名称重复")
|
||||
提供方 = {p.身份: p for p in self.提供方}
|
||||
if len(提供方) != len(self.提供方) or not 提供方:
|
||||
raise 配置版本错误("提供方必须非空且身份唯一")
|
||||
if any(p.凭据名称 not in {r.名称 for r in self.凭据} for p in self.提供方):
|
||||
raise 配置版本错误("提供方引用了未登记凭据")
|
||||
if any(p["provider"] not in 提供方 for p in self.角色配置.values()):
|
||||
raise 配置版本错误("角色引用了未登记提供方")
|
||||
if self.宿主 == "pi" and not all(
|
||||
isinstance(p, str) and Path(p).is_absolute() for p in (self.Node路径, self.Pi包目录)
|
||||
):
|
||||
raise 配置版本错误("Pi配置必须固定Node与包的绝对位置")
|
||||
object.__setattr__(self, "角色配置", json.loads(json.dumps(self.角色配置)))
|
||||
|
||||
def 冻结(self) -> dict:
|
||||
return json.loads(json.dumps(asdict(self), ensure_ascii=False))
|
||||
|
||||
@classmethod
|
||||
def 从快照(cls, 内容: dict) -> 运行配置内容:
|
||||
return cls(
|
||||
**{
|
||||
**内容,
|
||||
"凭据": tuple(凭据引用(**r) for r in 内容["凭据"]),
|
||||
"提供方": tuple(提供方配置(**p) for p in 内容["提供方"]),
|
||||
}
|
||||
)
|
||||
|
||||
|
||||
@dataclass(frozen=True, slots=True)
|
||||
class 配置验证证据:
|
||||
配置哈希: str
|
||||
角色策略版本: str
|
||||
资源发布身份: str
|
||||
执行用途: 用途
|
||||
证据引用: tuple[str, ...]
|
||||
模式: str
|
||||
|
||||
|
||||
class 配置验证器(Protocol):
|
||||
"""由装配登记的角色与资源验证实现;接入层不能自报验证成功。"""
|
||||
|
||||
@property
|
||||
def 身份(self) -> str: ...
|
||||
|
||||
def 验证(self, 内容: 运行配置内容, 执行用途: 用途) -> 配置验证证据: ...
|
||||
|
||||
|
||||
@dataclass(frozen=True, slots=True)
|
||||
class 配置快照:
|
||||
配置ID: str
|
||||
版本: str
|
||||
内容哈希: str
|
||||
内容: 运行配置内容
|
||||
|
||||
|
||||
class 配置版本管理:
|
||||
def __init__(self, 数据库: 数据库工厂, 验证器: 配置验证器 | None = None) -> None:
|
||||
self.数据库 = 数据库
|
||||
self.验证器 = 验证器
|
||||
self._schema = "evaluation" if 数据库.用途 is 用途.评测 else "public"
|
||||
|
||||
def _查询(self, 连: psycopg.Connection, 语句: LiteralString, 参数: tuple = ()) -> Any:
|
||||
return 连.cursor(row_factory=dict_row).execute(
|
||||
sql.SQL(语句).format(s=sql.Identifier(self._schema)), 参数
|
||||
)
|
||||
|
||||
@contextmanager
|
||||
def _事务(self) -> Iterator[psycopg.Connection]:
|
||||
with self.数据库.连接() as 连, 连.transaction():
|
||||
yield 连
|
||||
|
||||
def 保存草案(self, 配置ID: str, 版本: str, 内容: 运行配置内容) -> 配置快照:
|
||||
if not 配置ID or not 版本:
|
||||
raise 配置版本错误("配置身份与版本不能为空")
|
||||
冻结 = 运行配置内容.从快照(内容.冻结()).冻结()
|
||||
哈希 = 内容哈希(冻结)
|
||||
with self._事务() as 连:
|
||||
self._查询(
|
||||
连,
|
||||
"INSERT INTO {s}.muse_runtime_config_version "
|
||||
"(run_purpose,config_id,version,content,content_hash) VALUES (%s,%s,%s,%s,%s) "
|
||||
"ON CONFLICT DO NOTHING",
|
||||
(self.数据库.用途.value, 配置ID, 版本, Jsonb(冻结), 哈希),
|
||||
)
|
||||
已存 = self._版本(连, 配置ID, 版本)
|
||||
if 已存.内容哈希 != 哈希:
|
||||
raise 配置版本错误("同一配置版本不可覆盖,修改应建立新版本")
|
||||
return 已存
|
||||
|
||||
def 读取版本(self, 配置ID: str, 版本: str) -> 配置快照:
|
||||
with self._事务() as 连:
|
||||
return self._版本(连, 配置ID, 版本)
|
||||
|
||||
def _版本(self, 连: psycopg.Connection, 配置ID: str, 版本: str) -> 配置快照:
|
||||
行 = self._查询(
|
||||
连,
|
||||
"SELECT * FROM {s}.muse_runtime_config_version "
|
||||
"WHERE run_purpose=%s AND config_id=%s AND version=%s",
|
||||
(self.数据库.用途.value, 配置ID, 版本),
|
||||
).fetchone()
|
||||
if 行 is None:
|
||||
raise 配置版本错误("配置版本不存在或不属于当前用途")
|
||||
if 内容哈希(行["content"]) != 行["content_hash"]:
|
||||
raise 配置版本错误("配置内容与保存哈希不一致")
|
||||
return 配置快照(配置ID, 版本, 行["content_hash"], 运行配置内容.从快照(行["content"]))
|
||||
|
||||
def 验证版本(self, 配置ID: str, 版本: str) -> str:
|
||||
if self.验证器 is None:
|
||||
raise 配置版本错误("配置验证器尚未装配,不能自行声明验证通过")
|
||||
快照 = self.读取版本(配置ID, 版本)
|
||||
# 探针与资源核验由对应 owner 执行,验证期间不保持数据库事务。
|
||||
证据 = self.验证器.验证(快照.内容, self.数据库.用途)
|
||||
if (
|
||||
证据.配置哈希 != 快照.内容哈希
|
||||
or 证据.执行用途 is not self.数据库.用途
|
||||
or 证据.角色策略版本 != 快照.内容.角色策略版本
|
||||
or 证据.资源发布身份 != 快照.内容.资源发布身份
|
||||
):
|
||||
raise 配置版本错误("验证证据没有绑定当前配置、角色、资源与执行用途")
|
||||
if (
|
||||
not self.验证器.身份
|
||||
or not 证据.证据引用
|
||||
or 证据.模式 not in ("offline_contract", "runtime")
|
||||
):
|
||||
raise 配置版本错误("配置验证缺少明确模式与可回查证据")
|
||||
if self.数据库.用途 is 用途.生产 and 证据.模式 != "runtime":
|
||||
raise 配置版本错误("离线协议证据不能启用生产配置")
|
||||
身份 = str(uuid.uuid4())
|
||||
with self._事务() as 连:
|
||||
if self._版本(连, 配置ID, 版本).内容哈希 != 快照.内容哈希:
|
||||
raise 配置版本错误("验证期间配置发生变化")
|
||||
self._查询(
|
||||
连,
|
||||
"INSERT INTO {s}.muse_runtime_config_validation "
|
||||
"(receipt_id,run_purpose,config_id,version,content_hash,validator_id,"
|
||||
"evidence_refs) "
|
||||
"VALUES (%s,%s,%s,%s,%s,%s,%s)",
|
||||
(
|
||||
身份,
|
||||
self.数据库.用途.value,
|
||||
配置ID,
|
||||
版本,
|
||||
快照.内容哈希,
|
||||
self.验证器.身份,
|
||||
Jsonb({"模式": 证据.模式, "引用": list(证据.证据引用)}),
|
||||
),
|
||||
)
|
||||
return 身份
|
||||
|
||||
def _锁指针(self, 连: psycopg.Connection, 配置ID: str) -> dict:
|
||||
self._查询(
|
||||
连,
|
||||
"INSERT INTO {s}.muse_runtime_config_active (run_purpose,config_id) "
|
||||
"VALUES (%s,%s) ON CONFLICT DO NOTHING",
|
||||
(self.数据库.用途.value, 配置ID),
|
||||
)
|
||||
return self._查询(
|
||||
连,
|
||||
"SELECT * FROM {s}.muse_runtime_config_active "
|
||||
"WHERE run_purpose=%s AND config_id=%s FOR UPDATE",
|
||||
(self.数据库.用途.value, 配置ID),
|
||||
).fetchone()
|
||||
|
||||
def 启用(self, 配置ID: str, 版本: str, *, 验证回执: str, 批准引用: str, 预期代次: int) -> int:
|
||||
if not 批准引用:
|
||||
raise 配置版本错误("配置启用需要明确批准引用")
|
||||
with self._事务() as 连:
|
||||
当前 = self._锁指针(连, 配置ID)
|
||||
if (
|
||||
当前["version"] == 版本
|
||||
and str(当前["validation_receipt"]) == 验证回执
|
||||
and 当前["approval_ref"] == 批准引用
|
||||
):
|
||||
return 当前["generation"]
|
||||
if 当前["generation"] != 预期代次:
|
||||
raise 配置版本错误("配置启用指针已变化")
|
||||
快照 = self._版本(连, 配置ID, 版本)
|
||||
证据 = self._查询(
|
||||
连,
|
||||
"SELECT * FROM {s}.muse_runtime_config_validation "
|
||||
"WHERE receipt_id=%s AND run_purpose=%s AND config_id=%s AND version=%s "
|
||||
"AND content_hash=%s",
|
||||
(验证回执, self.数据库.用途.value, 配置ID, 版本, 快照.内容哈希),
|
||||
).fetchone()
|
||||
if 证据 is None:
|
||||
raise 配置版本错误("验证回执不属于本配置版本")
|
||||
新代次 = 当前["generation"] + 1
|
||||
self._查询(
|
||||
连,
|
||||
"UPDATE {s}.muse_runtime_config_active SET version=%s,validation_receipt=%s,"
|
||||
"approval_ref=%s,generation=%s WHERE run_purpose=%s AND config_id=%s",
|
||||
(版本, 验证回执, 批准引用, 新代次, self.数据库.用途.value, 配置ID),
|
||||
)
|
||||
return 新代次
|
||||
|
||||
def 停用(self, 配置ID: str, *, 预期代次: int) -> int:
|
||||
with self._事务() as 连:
|
||||
当前 = self._锁指针(连, 配置ID)
|
||||
if 当前["generation"] != 预期代次:
|
||||
raise 配置版本错误("配置停用指针已变化")
|
||||
新代次 = 当前["generation"] + 1
|
||||
self._查询(
|
||||
连,
|
||||
"UPDATE {s}.muse_runtime_config_active SET version=NULL,validation_receipt=NULL,"
|
||||
"approval_ref=NULL,generation=%s WHERE run_purpose=%s AND config_id=%s",
|
||||
(新代次, self.数据库.用途.value, 配置ID),
|
||||
)
|
||||
return 新代次
|
||||
|
||||
def 冻结到任务(self, 任务ID: str, 配置ID: str) -> 配置快照:
|
||||
with self._事务() as 连:
|
||||
任务 = 任务存储(连, self.数据库.用途).任务行(任务ID, 锁定=True)
|
||||
已绑定 = self._查询(
|
||||
连, "SELECT * FROM {s}.muse_task_config_binding WHERE task_id=%s", (任务ID,)
|
||||
).fetchone()
|
||||
if 已绑定:
|
||||
if 已绑定["config_id"] != 配置ID:
|
||||
raise 配置版本错误("任务已经绑定另一配置身份")
|
||||
return 配置快照(
|
||||
配置ID,
|
||||
已绑定["version"],
|
||||
已绑定["content_hash"],
|
||||
运行配置内容.从快照(已绑定["content"]),
|
||||
)
|
||||
当前 = self._锁指针(连, 配置ID)
|
||||
if 当前["version"] is None:
|
||||
raise 配置版本错误("配置尚未验证启用或已停用")
|
||||
快照 = self._版本(连, 配置ID, 当前["version"])
|
||||
if (快照.内容.角色策略版本, 快照.内容.资源发布身份) != (
|
||||
任务["frozen_input"]["角色策略版本"],
|
||||
任务["frozen_input"]["资源发布身份"],
|
||||
):
|
||||
raise 配置版本错误("运行配置与任务冻结角色或资源版本不一致")
|
||||
self._查询(
|
||||
连,
|
||||
"INSERT INTO {s}.muse_task_config_binding "
|
||||
"(task_id,config_id,version,content_hash,content,validation_receipt) "
|
||||
"VALUES (%s,%s,%s,%s,%s,%s)",
|
||||
(
|
||||
任务ID,
|
||||
配置ID,
|
||||
快照.版本,
|
||||
快照.内容哈希,
|
||||
Jsonb(快照.内容.冻结()),
|
||||
当前["validation_receipt"],
|
||||
),
|
||||
)
|
||||
return 快照
|
||||
|
||||
def 读取任务绑定(self, 任务ID: str) -> 配置快照:
|
||||
with self._事务() as 连:
|
||||
行 = self._查询(
|
||||
连, "SELECT * FROM {s}.muse_task_config_binding WHERE task_id=%s", (任务ID,)
|
||||
).fetchone()
|
||||
if 行 is None:
|
||||
raise 配置版本错误("任务尚未绑定验证启用的运行配置")
|
||||
if 内容哈希(行["content"]) != 行["content_hash"]:
|
||||
raise 配置版本错误("任务运行配置快照与保存哈希不一致")
|
||||
return 配置快照(
|
||||
行["config_id"],
|
||||
行["version"],
|
||||
行["content_hash"],
|
||||
运行配置内容.从快照(行["content"]),
|
||||
)
|
||||
|
||||
|
||||
__all__ = [
|
||||
"提供方配置",
|
||||
"凭据引用",
|
||||
"运行配置内容",
|
||||
"配置验证证据",
|
||||
"配置验证器",
|
||||
"配置快照",
|
||||
"配置版本管理",
|
||||
"配置版本错误",
|
||||
]
|
||||
556
src/muse/任务运行/预算管理.py
Normal file
556
src/muse/任务运行/预算管理.py
Normal file
@ -0,0 +1,556 @@
|
||||
"""版本化额度与逐调用预算账本;预留、取消和结算都在短事务中完成。"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import math
|
||||
from collections.abc import Iterator
|
||||
from contextlib import contextmanager
|
||||
from dataclasses import asdict, dataclass, replace
|
||||
from datetime import datetime, timedelta
|
||||
from decimal import ROUND_HALF_UP, Decimal, InvalidOperation
|
||||
from typing import Any, LiteralString
|
||||
from zoneinfo import ZoneInfo
|
||||
|
||||
import psycopg
|
||||
from psycopg import sql
|
||||
from psycopg.rows import dict_row
|
||||
from psycopg.types.json import Jsonb
|
||||
|
||||
from muse.任务运行.任务领取 import 校验领取
|
||||
from muse.任务运行.存储 import 任务存储
|
||||
from muse.任务运行.模型 import 内容哈希, 领取凭证
|
||||
from muse.共享.调用身份 import 用途
|
||||
from muse.共享.错误 import Muse错误
|
||||
from muse.基础设施.数据库.连接 import 数据库工厂
|
||||
|
||||
_金额精度 = Decimal("0.000001")
|
||||
_角色 = frozenset({"writer", "planner", "extractor", "detector", "judge"})
|
||||
|
||||
|
||||
class 预算错误(Muse错误):
|
||||
错误码 = "MUSE_BUDGET"
|
||||
|
||||
|
||||
class 预算不足(预算错误):
|
||||
错误码 = "MUSE_BUDGET_EXHAUSTED"
|
||||
|
||||
|
||||
class 预算状态冲突(预算错误):
|
||||
错误码 = "MUSE_BUDGET_STATE"
|
||||
|
||||
|
||||
def 金额(值: Decimal | str | int) -> Decimal:
|
||||
try:
|
||||
if isinstance(值, bool):
|
||||
raise ValueError
|
||||
结果 = Decimal(str(值))
|
||||
if not 结果.is_finite() or 结果 < 0:
|
||||
raise ValueError
|
||||
return 结果.quantize(_金额精度, rounding=ROUND_HALF_UP)
|
||||
except (InvalidOperation, ValueError):
|
||||
raise 预算错误("金额必须是明确的非负有限数") from None
|
||||
|
||||
|
||||
def _时区时间(值: datetime) -> datetime:
|
||||
if 值.tzinfo is None or 值.utcoffset() is None:
|
||||
raise 预算错误("预算时间必须包含时区")
|
||||
return 值
|
||||
|
||||
|
||||
@dataclass(frozen=True, slots=True)
|
||||
class 额度窗口:
|
||||
起点: datetime
|
||||
终点: datetime
|
||||
|
||||
|
||||
def 归属窗口(当前: datetime, 时区: str = "Asia/Shanghai") -> 额度窗口:
|
||||
"""沿用每日 00/05/10/15/20 边界;末窗到午夜,避免默改既有额度合同。"""
|
||||
本地 = _时区时间(当前).astimezone(ZoneInfo(时区))
|
||||
起点 = 本地.replace(hour=(本地.hour // 5) * 5, minute=0, second=0, microsecond=0)
|
||||
终点 = (
|
||||
起点 + timedelta(hours=5) if 起点.hour < 20 else (起点 + timedelta(days=1)).replace(hour=0)
|
||||
)
|
||||
return 额度窗口(起点, 终点)
|
||||
|
||||
|
||||
@dataclass(frozen=True, slots=True)
|
||||
class 额度策略:
|
||||
账户ID: str
|
||||
版本: str
|
||||
窗口金额上限: Decimal
|
||||
窗口调用上限: int
|
||||
时区: str = "Asia/Shanghai"
|
||||
|
||||
def __post_init__(self) -> None:
|
||||
object.__setattr__(self, "窗口金额上限", 金额(self.窗口金额上限))
|
||||
if not self.账户ID or not self.版本 or self.窗口金额上限 <= 0:
|
||||
raise 预算错误("额度策略需要身份、版本和正数金额上限")
|
||||
if type(self.窗口调用上限) is not int or self.窗口调用上限 < 1:
|
||||
raise 预算错误("窗口调用上限必须是正整数")
|
||||
ZoneInfo(self.时区)
|
||||
|
||||
def 冻结(self) -> dict:
|
||||
return {**asdict(self), "窗口金额上限": str(self.窗口金额上限)}
|
||||
|
||||
|
||||
@dataclass(frozen=True, slots=True)
|
||||
class 角色预算:
|
||||
角色: str
|
||||
计划次数: int
|
||||
最多次数: int
|
||||
单次上限: Decimal
|
||||
|
||||
def __post_init__(self) -> None:
|
||||
object.__setattr__(self, "单次上限", 金额(self.单次上限))
|
||||
if self.角色 not in _角色 or self.单次上限 <= 0:
|
||||
raise 预算错误("角色预算需要已登记角色和正数单次上限")
|
||||
if any(type(n) is not int or n < 1 for n in (self.计划次数, self.最多次数)):
|
||||
raise 预算错误("调用计划与安全容量必须为正整数")
|
||||
if self.计划次数 > self.最多次数:
|
||||
raise 预算错误("计划调用次数不能超过安全容量")
|
||||
|
||||
|
||||
@dataclass(frozen=True, slots=True)
|
||||
class 任务预算计划:
|
||||
总金额: Decimal
|
||||
角色: tuple[角色预算, ...]
|
||||
批准引用: str
|
||||
截止时间: datetime
|
||||
|
||||
def __post_init__(self) -> None:
|
||||
object.__setattr__(self, "总金额", 金额(self.总金额))
|
||||
_时区时间(self.截止时间)
|
||||
if not self.批准引用 or not self.角色 or len({r.角色 for r in self.角色}) != len(self.角色):
|
||||
raise 预算错误("任务预算需要批准引用和不重复的完整角色计划")
|
||||
if self.总金额 < self.最坏预留:
|
||||
raise 预算不足("任务总预算不能覆盖计划调用的最坏预留")
|
||||
|
||||
@property
|
||||
def 最坏预留(self) -> Decimal:
|
||||
return sum((r.单次上限 * r.计划次数 for r in self.角色), Decimal(0))
|
||||
|
||||
def 冻结(self) -> dict:
|
||||
return {
|
||||
"总金额": str(self.总金额),
|
||||
"批准引用": self.批准引用,
|
||||
"截止时间": self.截止时间.isoformat(),
|
||||
"角色": [{**asdict(r), "单次上限": str(r.单次上限)} for r in self.角色],
|
||||
}
|
||||
|
||||
|
||||
@dataclass(frozen=True, slots=True)
|
||||
class 预算预留:
|
||||
调用ID: str
|
||||
任务ID: str
|
||||
尝试ID: str
|
||||
角色: str
|
||||
状态: str
|
||||
预留金额: Decimal
|
||||
实际金额: Decimal | None
|
||||
到期时间: datetime
|
||||
结算回执: str | None
|
||||
超预算: bool
|
||||
允许外发: bool = False
|
||||
|
||||
|
||||
class 预算管理:
|
||||
def __init__(self, 数据库: 数据库工厂, 账户ID: str) -> None:
|
||||
self.数据库 = 数据库
|
||||
self.账户ID = 账户ID
|
||||
self._schema = "evaluation" if 数据库.用途 is 用途.评测 else "public"
|
||||
|
||||
def _查询(self, 连: psycopg.Connection, 语句: LiteralString, 参数: tuple = ()) -> Any:
|
||||
return 连.cursor(row_factory=dict_row).execute(
|
||||
sql.SQL(语句).format(s=sql.Identifier(self._schema)), 参数
|
||||
)
|
||||
|
||||
@contextmanager
|
||||
def _事务(self) -> Iterator[psycopg.Connection]:
|
||||
with self.数据库.连接() as 连, 连.transaction():
|
||||
yield 连
|
||||
|
||||
def _锁账户(self, 连: psycopg.Connection) -> dict:
|
||||
行 = self._查询(
|
||||
连,
|
||||
"SELECT * FROM {s}.muse_budget_account "
|
||||
"WHERE run_purpose=%s AND account_id=%s FOR UPDATE",
|
||||
(self.数据库.用途.value, self.账户ID),
|
||||
).fetchone()
|
||||
if 行 is None:
|
||||
raise 预算错误("额度账户尚未登记")
|
||||
return 行["policy"]
|
||||
|
||||
def 登记策略(self, 策略: 额度策略) -> None:
|
||||
if 策略.账户ID != self.账户ID:
|
||||
raise 预算错误("额度策略与账户不匹配")
|
||||
内容 = 策略.冻结()
|
||||
with self._事务() as 连:
|
||||
self._查询(
|
||||
连,
|
||||
"INSERT INTO {s}.muse_budget_account "
|
||||
"(run_purpose,account_id,policy_version,policy,policy_hash) VALUES (%s,%s,%s,"
|
||||
"%s,%s) "
|
||||
"ON CONFLICT DO NOTHING",
|
||||
(self.数据库.用途.value, self.账户ID, 策略.版本, Jsonb(内容), 内容哈希(内容)),
|
||||
)
|
||||
if self._锁账户(连) != 内容:
|
||||
raise 预算状态冲突("账户已有冻结额度策略;不能就地改变在途预算")
|
||||
|
||||
def 登记任务预算(self, 任务ID: str, 计划: 任务预算计划) -> None:
|
||||
内容 = 计划.冻结()
|
||||
with self._事务() as 连:
|
||||
任务存储(连, self.数据库.用途).任务行(任务ID, 锁定=True)
|
||||
self._锁账户(连)
|
||||
self._查询(
|
||||
连,
|
||||
"INSERT INTO {s}.muse_task_budget (task_id,account_id,plan,plan_hash) "
|
||||
"VALUES (%s,%s,%s,%s) ON CONFLICT DO NOTHING",
|
||||
(任务ID, self.账户ID, Jsonb(内容), 内容哈希(内容)),
|
||||
)
|
||||
既有 = self._任务预算(连, 任务ID)
|
||||
if 既有["plan_hash"] != 内容哈希(内容):
|
||||
raise 预算状态冲突("任务预算已冻结,重复登记不能改变批准金额与计划")
|
||||
|
||||
def _任务预算(self, 连: psycopg.Connection, 任务ID: str) -> dict:
|
||||
行 = self._查询(
|
||||
连,
|
||||
"SELECT * FROM {s}.muse_task_budget WHERE task_id=%s AND account_id=%s",
|
||||
(任务ID, self.账户ID),
|
||||
).fetchone()
|
||||
if 行 is None:
|
||||
raise 预算错误("任务尚未登记本账户的预算计划")
|
||||
return 行
|
||||
|
||||
def 回收过期(self) -> dict[str, int]:
|
||||
"""未发送预留可释放;已发送过期只能转未知,不能自动抹掉成本。"""
|
||||
with self._事务() as 连:
|
||||
self._锁账户(连)
|
||||
条件 = (self.数据库.用途.value, self.账户ID)
|
||||
释放 = self._查询(
|
||||
连,
|
||||
"UPDATE {s}.muse_budget_reservation SET state='released',reason='expired' "
|
||||
"WHERE run_purpose=%s AND account_id=%s AND state='reserved' "
|
||||
"AND expires_at<=clock_timestamp()",
|
||||
条件,
|
||||
).rowcount
|
||||
未知 = self._查询(
|
||||
连,
|
||||
"UPDATE {s}.muse_budget_reservation SET state='unknown',reason='expired' "
|
||||
"WHERE run_purpose=%s AND account_id=%s AND state='in_flight' "
|
||||
"AND expires_at<=clock_timestamp()",
|
||||
条件,
|
||||
).rowcount
|
||||
return {"释放": 释放, "未知": 未知}
|
||||
|
||||
def _调用行(self, 连: psycopg.Connection, 调用ID: str) -> dict | None:
|
||||
return self._查询(
|
||||
连,
|
||||
"SELECT * FROM {s}.muse_budget_reservation "
|
||||
"WHERE run_purpose=%s AND account_id=%s AND call_id=%s",
|
||||
(self.数据库.用途.value, self.账户ID, 调用ID),
|
||||
).fetchone()
|
||||
|
||||
def _所有调用(self, 连: psycopg.Connection) -> list[dict]:
|
||||
return self._查询(
|
||||
连,
|
||||
"SELECT * FROM {s}.muse_budget_reservation "
|
||||
"WHERE run_purpose=%s AND account_id=%s AND state<>'released'",
|
||||
(self.数据库.用途.value, self.账户ID),
|
||||
).fetchall()
|
||||
|
||||
def 预留(self, 领取: 领取凭证, 调用ID: str, 角色: str, *, 有效秒: float = 60) -> 预算预留:
|
||||
if not 调用ID or not math.isfinite(有效秒) or 有效秒 <= 0:
|
||||
raise 预算错误("调用身份和正数预留期限必须提供")
|
||||
self.回收过期()
|
||||
with self._事务() as 连:
|
||||
校验领取(任务存储(连, self.数据库.用途), 领取)
|
||||
策略 = self._锁账户(连)
|
||||
请求哈希 = 内容哈希(
|
||||
{
|
||||
"task": 领取.任务ID,
|
||||
"attempt": 领取.尝试ID,
|
||||
"role": 角色,
|
||||
"account": self.账户ID,
|
||||
"ttl": 有效秒,
|
||||
}
|
||||
)
|
||||
既有 = self._调用行(连, 调用ID)
|
||||
if 既有:
|
||||
if 既有["request_hash"] != 请求哈希:
|
||||
raise 预算状态冲突("同一调用身份不能重复绑定不同预算请求")
|
||||
return self._预留(既有)
|
||||
当前 = self._查询(连, "SELECT clock_timestamp() AS current_time").fetchone()[
|
||||
"current_time"
|
||||
]
|
||||
窗口 = 归属窗口(当前, 策略["时区"])
|
||||
登记 = self._任务预算(连, 领取.任务ID)
|
||||
计划 = 登记["plan"]
|
||||
if 登记["stopped"] or datetime.fromisoformat(计划["截止时间"]) <= 当前:
|
||||
raise 预算不足("任务预算已停止或超过总期限")
|
||||
所有 = self._所有调用(连)
|
||||
if any(r["state"] == "unknown" for r in 所有):
|
||||
raise 预算不足("账户存在未知成本,必须先对账")
|
||||
本窗 = [r for r in 所有 if r["window_start"] == 窗口.起点]
|
||||
已占 = sum((self._占用(r) for r in 本窗), Decimal(0))
|
||||
角色表 = {r["角色"]: r for r in 计划["角色"]}
|
||||
if 角色 not in 角色表:
|
||||
raise 预算错误("当前角色未纳入已批准的任务调用计划")
|
||||
每次 = Decimal(角色表[角色]["单次上限"])
|
||||
if len(本窗) >= 策略["窗口调用上限"] or 已占 + 每次 > Decimal(策略["窗口金额上限"]):
|
||||
raise 预算不足(
|
||||
"当前窗口的金额或调用次数不足", 上下文={"下个窗口": 窗口.终点.isoformat()}
|
||||
)
|
||||
本任务 = [r for r in 所有 if str(r["task_id"]) == 领取.任务ID]
|
||||
用量 = {role: sum(r["role_id"] == role for r in 本任务) for role in 角色表}
|
||||
目标 = 角色表[角色]
|
||||
if 用量[角色] >= min(目标["计划次数"], 目标["最多次数"]):
|
||||
raise 预算不足("当前角色的计划调用次数已耗尽")
|
||||
剩余 = sum(
|
||||
(
|
||||
Decimal(r["单次上限"]) * max(r["计划次数"] - 用量[role], 0)
|
||||
for role, r in 角色表.items()
|
||||
),
|
||||
Decimal(0),
|
||||
)
|
||||
if sum((self._占用(r) for r in 本任务), Decimal(0)) + 剩余 > Decimal(计划["总金额"]):
|
||||
raise 预算不足("实际成本与剩余计划预留超过任务总预算")
|
||||
到期 = min(
|
||||
当前 + timedelta(seconds=有效秒),
|
||||
窗口.终点,
|
||||
datetime.fromisoformat(计划["截止时间"]),
|
||||
)
|
||||
行 = self._查询(
|
||||
连,
|
||||
"INSERT INTO {s}.muse_budget_reservation "
|
||||
"(run_purpose,call_id,account_id,task_id,attempt_id,role_id,request_hash,state,"
|
||||
"reserved_amount,window_start,window_end,expires_at) "
|
||||
"VALUES (%s,%s,%s,%s,%s,%s,%s,'reserved',%s,%s,%s,%s) RETURNING *",
|
||||
(
|
||||
self.数据库.用途.value,
|
||||
调用ID,
|
||||
self.账户ID,
|
||||
领取.任务ID,
|
||||
领取.尝试ID,
|
||||
角色,
|
||||
请求哈希,
|
||||
每次,
|
||||
窗口.起点,
|
||||
窗口.终点,
|
||||
到期,
|
||||
),
|
||||
).fetchone()
|
||||
return self._预留(行)
|
||||
|
||||
def 标记已发送(
|
||||
self, 领取: 领取凭证, 调用ID: str, *, 最长秒: float, 登记尝试: bool = False
|
||||
) -> 预算预留:
|
||||
if not math.isfinite(最长秒) or 最长秒 <= 0:
|
||||
raise 预算错误("调用总期限必须大于零")
|
||||
self.回收过期()
|
||||
with self._事务() as 连:
|
||||
存储 = 任务存储(连, self.数据库.用途)
|
||||
_, 尝试 = 校验领取(存储, 领取)
|
||||
self._锁账户(连)
|
||||
行 = self._要求调用(连, 调用ID)
|
||||
if (str(行["task_id"]), str(行["attempt_id"])) != (领取.任务ID, 领取.尝试ID):
|
||||
raise 预算状态冲突("调用预留不属于当前尝试")
|
||||
if 行["state"] == "in_flight":
|
||||
return self._预留(行)
|
||||
当前 = self._查询(连, "SELECT clock_timestamp() AS current_time").fetchone()[
|
||||
"current_time"
|
||||
]
|
||||
if 行["state"] != "reserved" or 行["expires_at"] <= 当前:
|
||||
raise 预算状态冲突("只有有效且尚未发送的预留可以外发")
|
||||
登记 = self._任务预算(连, 领取.任务ID)
|
||||
if 登记["stopped"]:
|
||||
raise 预算不足("任务预算已经停止")
|
||||
if any(r["state"] == "unknown" for r in self._所有调用(连)):
|
||||
raise 预算不足("账户存在未知成本,必须先对账")
|
||||
到期 = min(
|
||||
当前 + timedelta(seconds=最长秒), datetime.fromisoformat(登记["plan"]["截止时间"])
|
||||
)
|
||||
if 到期 <= 当前:
|
||||
raise 预算不足("调用已经超过任务总期限")
|
||||
if 登记尝试:
|
||||
if 尝试["call_state"] not in {"not_sent", "saved"}:
|
||||
raise 预算状态冲突("本次尝试的上一调用仍未完成证据对账")
|
||||
存储.登记调用(领取, 调用ID)
|
||||
self._查询(
|
||||
连,
|
||||
"UPDATE {s}.muse_budget_reservation SET state='in_flight',sent_at=%s,expires_at=%s "
|
||||
"WHERE run_purpose=%s AND call_id=%s",
|
||||
(当前, 到期, self.数据库.用途.value, 调用ID),
|
||||
)
|
||||
return replace(self._预留(self._要求调用(连, 调用ID)), 允许外发=True)
|
||||
|
||||
def 结算(self, 调用ID: str, 实际金额: Decimal | None, *, 回执ID: str) -> 预算预留:
|
||||
if not 回执ID:
|
||||
raise 预算错误("结算需要实际回执身份")
|
||||
成本 = None if 实际金额 is None else 金额(实际金额)
|
||||
with self._事务() as 连:
|
||||
self._锁账户(连)
|
||||
行 = self._要求调用(连, 调用ID)
|
||||
if 行["state"] == "settled":
|
||||
if 行["receipt_id"] != 回执ID or 行["actual_amount"] != 成本:
|
||||
raise 预算状态冲突("调用已经结算,不能改写成本或回执")
|
||||
return self._预留(行)
|
||||
if 行["state"] not in ("in_flight", "unknown"):
|
||||
raise 预算状态冲突("未发送的调用不能记作实际成本")
|
||||
超额 = 成本 is not None and 成本 > 行["reserved_amount"]
|
||||
self._查询(
|
||||
连,
|
||||
"UPDATE {s}.muse_budget_reservation SET state=%s,actual_amount=%s,"
|
||||
"receipt_id=%s,over_budget=%s WHERE run_purpose=%s AND call_id=%s",
|
||||
(
|
||||
"unknown" if 成本 is None else "settled",
|
||||
成本,
|
||||
回执ID,
|
||||
超额,
|
||||
self.数据库.用途.value,
|
||||
调用ID,
|
||||
),
|
||||
)
|
||||
if 超额:
|
||||
self._查询(
|
||||
连,
|
||||
"UPDATE {s}.muse_task_budget SET stopped=true WHERE task_id=%s",
|
||||
(行["task_id"],),
|
||||
)
|
||||
return self._预留(self._要求调用(连, 调用ID))
|
||||
|
||||
def 取消预留(self, 调用ID: str) -> 预算预留:
|
||||
with self._事务() as 连:
|
||||
self._锁账户(连)
|
||||
行 = self._要求调用(连, 调用ID)
|
||||
self._取消行(连, 行)
|
||||
return self._预留(self._要求调用(连, 调用ID))
|
||||
|
||||
def 关闭任务预算(self, 任务ID: str) -> None:
|
||||
with self._事务() as 连:
|
||||
存储 = 任务存储(连, self.数据库.用途)
|
||||
存储.任务行(任务ID, 锁定=True)
|
||||
停止任务预算(存储, 任务ID)
|
||||
|
||||
def _取消行(self, 连: psycopg.Connection, 行: dict) -> None:
|
||||
if 行["state"] in ("reserved", "in_flight"):
|
||||
self._查询(
|
||||
连,
|
||||
"UPDATE {s}.muse_budget_reservation SET state=%s,reason='cancelled' "
|
||||
"WHERE run_purpose=%s AND call_id=%s",
|
||||
(
|
||||
"released" if 行["state"] == "reserved" else "unknown",
|
||||
self.数据库.用途.value,
|
||||
行["call_id"],
|
||||
),
|
||||
)
|
||||
|
||||
def 读取(self, 调用ID: str) -> 预算预留:
|
||||
with self._事务() as 连:
|
||||
return self._预留(self._要求调用(连, 调用ID))
|
||||
|
||||
def 窗口余额(self) -> dict:
|
||||
self.回收过期()
|
||||
with self._事务() as 连:
|
||||
策略 = self._锁账户(连)
|
||||
当前 = self._查询(连, "SELECT clock_timestamp() AS current_time").fetchone()[
|
||||
"current_time"
|
||||
]
|
||||
窗口 = 归属窗口(当前, 策略["时区"])
|
||||
所有 = self._所有调用(连)
|
||||
本窗 = [r for r in 所有 if r["window_start"] == 窗口.起点]
|
||||
已知 = sum((r["actual_amount"] for r in 本窗 if r["state"] == "settled"), Decimal(0))
|
||||
在途 = sum((r["reserved_amount"] for r in 本窗 if r["state"] != "settled"), Decimal(0))
|
||||
return {
|
||||
"已知成本": 已知,
|
||||
"在途预留": 在途,
|
||||
"已占次数": len(本窗),
|
||||
"未知调用数": sum(r["state"] == "unknown" for r in 所有),
|
||||
"可用金额": max(Decimal(0), Decimal(策略["窗口金额上限"]) - 已知 - 在途),
|
||||
"窗口起点": 窗口.起点,
|
||||
"窗口终点": 窗口.终点,
|
||||
}
|
||||
|
||||
def _要求调用(self, 连: psycopg.Connection, 调用ID: str) -> dict:
|
||||
行 = self._调用行(连, 调用ID)
|
||||
if 行 is None:
|
||||
raise 预算状态冲突("调用预算不存在或不属于本账户用途")
|
||||
return 行
|
||||
|
||||
@staticmethod
|
||||
def _占用(行: dict) -> Decimal:
|
||||
return 行["actual_amount"] if 行["state"] == "settled" else 行["reserved_amount"]
|
||||
|
||||
@staticmethod
|
||||
def _预留(行: dict) -> 预算预留:
|
||||
return 预算预留(
|
||||
行["call_id"],
|
||||
str(行["task_id"]),
|
||||
str(行["attempt_id"]),
|
||||
行["role_id"],
|
||||
行["state"],
|
||||
行["reserved_amount"],
|
||||
行["actual_amount"],
|
||||
行["expires_at"],
|
||||
行["receipt_id"],
|
||||
行["over_budget"],
|
||||
)
|
||||
|
||||
|
||||
def 恢复任务预算(存储: 任务存储, 任务ID: str) -> None:
|
||||
"""作者恢复在原批准范围内重新开放剩余额度,不重置已使用次数和费用。"""
|
||||
预算 = 存储.查询(
|
||||
"SELECT account_id,plan FROM {s}.muse_task_budget WHERE task_id=%s", (任务ID,)
|
||||
).fetchone()
|
||||
if 预算 is None:
|
||||
return
|
||||
存储.查询(
|
||||
"SELECT account_id FROM {s}.muse_budget_account "
|
||||
"WHERE run_purpose=%s AND account_id=%s FOR UPDATE",
|
||||
(存储.用途.value, 预算["account_id"]),
|
||||
)
|
||||
当前 = 存储.查询("SELECT clock_timestamp() AS now").fetchone()["now"]
|
||||
if datetime.fromisoformat(预算["plan"]["截止时间"]) <= 当前:
|
||||
raise 预算不足("原任务预算已经截止")
|
||||
if 存储.查询(
|
||||
"SELECT call_id FROM {s}.muse_budget_reservation WHERE task_id=%s "
|
||||
"AND state IN ('unknown','in_flight') LIMIT 1",
|
||||
(任务ID,),
|
||||
).fetchone():
|
||||
raise 预算不足("恢复前必须对账所有在途调用")
|
||||
存储.查询("UPDATE {s}.muse_task_budget SET stopped=false WHERE task_id=%s", (任务ID,))
|
||||
|
||||
|
||||
def 停止任务预算(存储: 任务存储, 任务ID: str) -> None:
|
||||
"""与任务终止共用短事务;按任务、账户顺序锁定,不丢失已发出的成本。"""
|
||||
预算 = 存储.查询(
|
||||
"SELECT account_id FROM {s}.muse_task_budget WHERE task_id=%s", (任务ID,)
|
||||
).fetchone()
|
||||
if 预算 is None:
|
||||
return
|
||||
存储.查询(
|
||||
"SELECT account_id FROM {s}.muse_budget_account "
|
||||
"WHERE run_purpose=%s AND account_id=%s FOR UPDATE",
|
||||
(存储.用途.value, 预算["account_id"]),
|
||||
)
|
||||
存储.查询("UPDATE {s}.muse_task_budget SET stopped=true WHERE task_id=%s", (任务ID,))
|
||||
存储.查询(
|
||||
"UPDATE {s}.muse_budget_reservation SET state=CASE WHEN state='reserved' "
|
||||
"THEN 'released' ELSE 'unknown' END,reason='task_stopped' "
|
||||
"WHERE task_id=%s AND state IN ('reserved','in_flight')",
|
||||
(任务ID,),
|
||||
)
|
||||
|
||||
|
||||
__all__ = [
|
||||
"额度策略",
|
||||
"额度窗口",
|
||||
"归属窗口",
|
||||
"角色预算",
|
||||
"任务预算计划",
|
||||
"预算预留",
|
||||
"预算管理",
|
||||
"预算错误",
|
||||
"预算不足",
|
||||
"预算状态冲突",
|
||||
"金额",
|
||||
]
|
||||
124
src/muse/基础设施/宿主/Pi.py
Normal file
124
src/muse/基础设施/宿主/Pi.py
Normal file
@ -0,0 +1,124 @@
|
||||
"""固定Pi SDK的受控子进程;所有模型与工具请求回到S02,不提供直接外发凭据。"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import asyncio
|
||||
import json
|
||||
import os
|
||||
import subprocess
|
||||
from collections.abc import Awaitable, Callable
|
||||
from dataclasses import dataclass
|
||||
from pathlib import Path
|
||||
from tempfile import TemporaryDirectory
|
||||
|
||||
from muse.任务运行.接口 import 工具请求, 模型协议错误, 模型结果, 模型请求
|
||||
from muse.基础设施.宿主.Pi事件归一 import 归一事件, 检查会话结束
|
||||
|
||||
|
||||
@dataclass(frozen=True)
|
||||
class Pi宿主:
|
||||
Node路径: Path
|
||||
包目录: Path
|
||||
固定版本: str
|
||||
总期限秒: float = 120
|
||||
|
||||
def _预检(self) -> None:
|
||||
try:
|
||||
包 = json.loads((self.包目录 / "package.json").read_text())
|
||||
if 包["name"] != "@earendil-works/pi-coding-agent" or 包["version"] != self.固定版本:
|
||||
raise 模型协议错误("Pi安装身份或版本与固定配置不一致")
|
||||
版本 = subprocess.run(
|
||||
[str(self.Node路径), "--version"],
|
||||
check=True,
|
||||
capture_output=True,
|
||||
text=True,
|
||||
timeout=5,
|
||||
).stdout.strip()
|
||||
if tuple(map(int, 版本.removeprefix("v").split("."))) < (22, 19, 0):
|
||||
raise 模型协议错误("Pi桥需要Node 22.19.0及以上")
|
||||
except (OSError, ValueError, KeyError, subprocess.SubprocessError):
|
||||
raise 模型协议错误("Pi或Node安装不可用,不能自动降为无工具执行") from None
|
||||
|
||||
async def 运行(
|
||||
self,
|
||||
请求: 模型请求,
|
||||
模型回合: Callable[[], Awaitable[模型结果]],
|
||||
工具回合: Callable[[工具请求], Awaitable[str]],
|
||||
) -> 模型结果:
|
||||
self._预检()
|
||||
桥 = Path(__file__).with_name("Pi工具桥.ts")
|
||||
with TemporaryDirectory(prefix="muse-pi-session-") as 暂存:
|
||||
进程 = await asyncio.create_subprocess_exec(
|
||||
str(self.Node路径),
|
||||
str(桥),
|
||||
cwd=暂存,
|
||||
env={
|
||||
**{k: os.environ[k] for k in ("PATH", "LANG", "TMPDIR") if k in os.environ},
|
||||
"PI_OFFLINE": "1",
|
||||
"PI_TELEMETRY": "0",
|
||||
},
|
||||
stdin=asyncio.subprocess.PIPE,
|
||||
stdout=asyncio.subprocess.PIPE,
|
||||
stderr=asyncio.subprocess.DEVNULL,
|
||||
limit=16 * 1024 * 1024,
|
||||
)
|
||||
assert 进程.stdin is not None and 进程.stdout is not None
|
||||
|
||||
async def 发送(数据: dict) -> None:
|
||||
assert 进程.stdin is not None
|
||||
进程.stdin.write((json.dumps(数据, ensure_ascii=False) + "\n").encode())
|
||||
await 进程.stdin.drain()
|
||||
|
||||
事件, 最后, 已结束 = [], None, False
|
||||
try:
|
||||
async with asyncio.timeout(self.总期限秒):
|
||||
await 发送(
|
||||
{
|
||||
"type": "initialize",
|
||||
"package": str(self.包目录.resolve()),
|
||||
"system": 请求.系统提示,
|
||||
"input": 请求.用户输入,
|
||||
"model": 请求.model,
|
||||
"max_tokens": 请求.最大输出token,
|
||||
"tools": 请求.工具,
|
||||
}
|
||||
)
|
||||
while 行 := await 进程.stdout.readline():
|
||||
数据 = json.loads(行)
|
||||
if 数据["type"] == "event":
|
||||
归一事件(数据["event"])
|
||||
事件.append(数据["event"])
|
||||
elif 数据["type"] == "request":
|
||||
if 数据["kind"] == "model":
|
||||
最后 = await 模型回合()
|
||||
值 = {
|
||||
"state": 最后.状态,
|
||||
"text": 最后.文本,
|
||||
"tools": [
|
||||
{"id": t.调用ID, "name": t.名称, "arguments": t.参数}
|
||||
for t in 最后.工具调用
|
||||
],
|
||||
}
|
||||
elif 数据["kind"] == "tool":
|
||||
参数 = 数据["arguments"]
|
||||
文本 = await 工具回合(
|
||||
工具请求(参数["id"], 参数["name"], 参数["arguments"])
|
||||
)
|
||||
值 = {"text": 文本}
|
||||
else:
|
||||
raise 模型协议错误("Pi桥请求未登记能力")
|
||||
await 发送({"type": "reply", "id": 数据["id"], "ok": True, "value": 值})
|
||||
elif 数据["type"] == "finished":
|
||||
已结束 = True
|
||||
else:
|
||||
raise 模型协议错误("Pi桥未完成角色循环")
|
||||
退出码 = await 进程.wait()
|
||||
if 最后 is None or not 已结束:
|
||||
raise 模型协议错误("Pi没有交回最终模型回合")
|
||||
return 检查会话结束(事件, 退出码, 最后)
|
||||
except (TimeoutError, ValueError):
|
||||
raise 模型协议错误("Pi角色循环超时或协议帧无效") from None
|
||||
finally:
|
||||
if 进程.returncode is None:
|
||||
进程.kill()
|
||||
await 进程.wait()
|
||||
57
src/muse/基础设施/宿主/Pi事件归一.py
Normal file
57
src/muse/基础设施/宿主/Pi事件归一.py
Normal file
@ -0,0 +1,57 @@
|
||||
"""纯归一Pi通知;费用与模型身份取自S02回合,不以Pi会话结束代替业务通过。"""
|
||||
|
||||
from collections.abc import Iterable
|
||||
|
||||
from muse.任务运行.接口 import 模型协议错误, 模型结果
|
||||
|
||||
|
||||
def 归一事件(事件: dict) -> tuple[dict, ...]:
|
||||
类型 = 事件.get("type")
|
||||
if 类型 in {"auto_retry_start", "compaction_start", "summarization_retry_scheduled"}:
|
||||
raise 模型协议错误("Pi发起了未登记的自动续跑或压缩")
|
||||
if 类型 == "extension_error":
|
||||
raise 模型协议错误("Pi受控扩展执行失败")
|
||||
if 类型 == "agent_end" and 事件.get("willRetry"):
|
||||
raise 模型协议错误("Pi请求了未登记的自动重试")
|
||||
if 类型 in {"tool_execution_start", "tool_execution_end"}:
|
||||
if not 事件.get("toolCallId") or not 事件.get("toolName"):
|
||||
raise 模型协议错误("Pi工具事件缺少调用身份")
|
||||
标准 = {
|
||||
"type": "tool.started" if 类型 == "tool_execution_start" else "tool.completed",
|
||||
"call_id": 事件["toolCallId"],
|
||||
"tool": 事件["toolName"],
|
||||
}
|
||||
if 类型 == "tool_execution_end":
|
||||
标准["failed"] = bool(事件.get("isError"))
|
||||
return (标准,)
|
||||
if 类型 == "message_update":
|
||||
片段 = 事件.get("assistantMessageEvent", {})
|
||||
if 片段.get("type") == "text_delta":
|
||||
return ({"type": "model.text_delta", "text": 片段.get("delta", "")},)
|
||||
if 类型 == "agent_start":
|
||||
return ({"type": "host.started"},)
|
||||
if 类型 == "agent_settled":
|
||||
return ({"type": "host.settled"},)
|
||||
return ()
|
||||
|
||||
|
||||
def 检查会话结束(事件序列: Iterable[dict], 退出码: int, 最终回合: 模型结果) -> 模型结果:
|
||||
结束 = False
|
||||
最后消息 = None
|
||||
for 事件 in 事件序列:
|
||||
归一事件(事件)
|
||||
if 事件.get("type") == "message_end" and 事件.get("message", {}).get("role") == "assistant":
|
||||
最后消息 = 事件["message"]
|
||||
elif 事件.get("type") == "agent_end":
|
||||
候选 = [m for m in 事件.get("messages", []) if m.get("role") == "assistant"]
|
||||
if 候选:
|
||||
最后消息 = 候选[-1]
|
||||
elif 事件.get("type") == "agent_settled":
|
||||
结束 = True
|
||||
if 退出码 != 0 or not 结束 or 最后消息 is None or 最后消息.get("stopReason") != "stop":
|
||||
raise 模型协议错误("Pi未给出成功结束与完整最终消息")
|
||||
文本 = "".join(v["text"] for v in 最后消息.get("content", []) if v.get("type") == "text")
|
||||
if 最终回合.状态 != "completed" or 最终回合.工具调用 or 文本 != 最终回合.文本:
|
||||
raise 模型协议错误("Pi最终消息与已核对模型回合不一致")
|
||||
# 不使用SDK自报的模型名或费用重新结算,也不重复计算agent_end中的消息。
|
||||
return 最终回合
|
||||
131
src/muse/基础设施/宿主/Pi工具桥.ts
Normal file
131
src/muse/基础设施/宿主/Pi工具桥.ts
Normal file
@ -0,0 +1,131 @@
|
||||
// 仅连接父进程登记的S02回合与只读工具;提供方凭据和数据库连接不进入Pi。
|
||||
import { createInterface } from "node:readline";
|
||||
import { pathToFileURL } from "node:url";
|
||||
import { join } from "node:path";
|
||||
|
||||
type Packet = Record<string, any>;
|
||||
const lines = createInterface({ input: process.stdin });
|
||||
const pending = new Map<number, (value: Packet) => void>();
|
||||
let sequence = 0;
|
||||
let initialize: (value: Packet) => void;
|
||||
const initialized = new Promise<Packet>((resolve) => { initialize = resolve; });
|
||||
const send = (value: Packet) => process.stdout.write(JSON.stringify(value) + "\n");
|
||||
lines.on("line", (line) => {
|
||||
const packet = JSON.parse(line);
|
||||
if (packet.type === "initialize") initialize(packet);
|
||||
else {
|
||||
const resolve = pending.get(packet.id);
|
||||
if (!resolve) throw new Error("未知父进程响应");
|
||||
pending.delete(packet.id);
|
||||
resolve(packet);
|
||||
}
|
||||
});
|
||||
async function call(kind: string, arguments_: Packet = {}): Promise<Packet> {
|
||||
const id = ++sequence;
|
||||
const result = new Promise<Packet>((resolve) => pending.set(id, resolve));
|
||||
send({ type: "request", id, kind, arguments: arguments_ });
|
||||
const response = await result;
|
||||
if (!response.ok) throw new Error("S02拒绝本次调用");
|
||||
return response.value;
|
||||
}
|
||||
|
||||
const config = await initialized;
|
||||
const pkg = config.package;
|
||||
const { createAgentSession, DefaultResourceLoader, ModelRuntime, SessionManager, SettingsManager }
|
||||
= await import(pathToFileURL(join(pkg, "dist/index.js")).href);
|
||||
const { createAssistantMessageEventStream }
|
||||
= await import(pathToFileURL(join(pkg, "node_modules/@earendil-works/pi-ai/dist/index.js")).href);
|
||||
let session: any;
|
||||
try {
|
||||
const cwd = process.cwd();
|
||||
const runtime = await ModelRuntime.create({
|
||||
authPath: join(cwd, "auth.json"), modelsPath: null,
|
||||
modelsStorePath: join(cwd, "models-cache.json"), allowModelNetwork: false, refreshOnCreate: false,
|
||||
});
|
||||
// 这是进程内回调提供方,不发HTTP;SDK费用不作为S02计费事实。
|
||||
runtime.registerProvider("muse-governed", {
|
||||
api: "anthropic-messages", apiKey: "local-callback", baseUrl: "http://127.0.0.1:1",
|
||||
models: [{
|
||||
id: config.model, name: "Muse受控角色", reasoning: false, input: ["text"],
|
||||
contextWindow: 1000000, maxTokens: config.max_tokens,
|
||||
cost: { input: 0, output: 0, cacheRead: 0, cacheWrite: 0 },
|
||||
}],
|
||||
streamSimple(model: any, context: any) {
|
||||
if (context.systemPrompt !== config.system) throw new Error("Pi改变了固定系统提示");
|
||||
const stream = createAssistantMessageEventStream();
|
||||
void (async () => {
|
||||
try {
|
||||
const result = await call("model");
|
||||
const content: Packet[] = result.text ? [{ type: "text", text: result.text }] : [];
|
||||
content.push(...result.tools.map((t: any) => ({
|
||||
type: "toolCall", id: t.id, name: t.name, arguments: t.arguments,
|
||||
})));
|
||||
const message = {
|
||||
role: "assistant", content, api: model.api, provider: model.provider, model: model.id,
|
||||
usage: { input: 0, output: 0, cacheRead: 0, cacheWrite: 0, totalTokens: 0,
|
||||
cost: { input: 0, output: 0, cacheRead: 0, cacheWrite: 0, total: 0 } },
|
||||
stopReason: result.state === "tool_calls" ? "toolUse" : "stop", timestamp: Date.now(),
|
||||
};
|
||||
stream.push({ type: "start", partial: message });
|
||||
stream.push({ type: "done", reason: message.stopReason, message });
|
||||
stream.end();
|
||||
} catch {
|
||||
// 错误文字不复制模型输入、凭据或SDK堆栈到普通输出。
|
||||
send({ type: "fatal", code: "CALLBACK_FAILED" });
|
||||
process.exitCode = 1;
|
||||
lines.close();
|
||||
stream.end();
|
||||
}
|
||||
})();
|
||||
return stream;
|
||||
},
|
||||
});
|
||||
const settings = SettingsManager.inMemory({
|
||||
compaction: { enabled: false }, retry: { enabled: false, provider: { maxRetries: 0 } },
|
||||
defaultTools: [],
|
||||
});
|
||||
const loader = new DefaultResourceLoader({
|
||||
cwd, agentDir: cwd, settingsManager: settings, noExtensions: true, noSkills: true,
|
||||
noPromptTemplates: true, noThemes: true, noContextFiles: true,
|
||||
systemPromptOverride: () => config.system,
|
||||
extensionFactories: [(pi: any) => {
|
||||
pi.on("before_agent_start", () => ({ systemPrompt: config.system }));
|
||||
}],
|
||||
});
|
||||
await loader.reload();
|
||||
({ session } = await createAgentSession({
|
||||
cwd, agentDir: cwd, modelRuntime: runtime,
|
||||
model: runtime.getModel("muse-governed", config.model), thinkingLevel: "off",
|
||||
tools: config.tools.map((t: any) => t.name),
|
||||
customTools: config.tools.map((t: any) => ({
|
||||
name: t.name, label: t.description, description: t.description, parameters: t.parameters,
|
||||
execute: async (id: string, args: Packet) => {
|
||||
const result = await call("tool", { id, name: t.name, arguments: args });
|
||||
return { content: [{ type: "text", text: result.text }], details: { governed: true } };
|
||||
},
|
||||
})),
|
||||
resourceLoader: loader, settingsManager: settings, sessionManager: SessionManager.inMemory(),
|
||||
}));
|
||||
session.subscribe((event: Packet) => {
|
||||
const allowed = [
|
||||
"agent_start", "agent_settled", "auto_retry_start", "compaction_start",
|
||||
"summarization_retry_scheduled", "extension_error", "tool_execution_start", "tool_execution_end",
|
||||
];
|
||||
if (allowed.includes(event.type)) {
|
||||
send({ type: "event", event: { type: event.type, toolCallId: event.toolCallId,
|
||||
toolName: event.toolName, isError: event.isError } });
|
||||
} else if (event.type === "agent_end") {
|
||||
send({ type: "event", event: { type: event.type, willRetry: event.willRetry } });
|
||||
} else if (event.type === "message_end" && event.message?.role === "assistant") {
|
||||
send({ type: "event", event: { type: event.type, message: event.message } });
|
||||
}
|
||||
});
|
||||
await session.prompt(config.input);
|
||||
send({ type: "finished" });
|
||||
} catch {
|
||||
send({ type: "fatal", code: "PI_SESSION_FAILED" });
|
||||
process.exitCode = 1;
|
||||
} finally {
|
||||
session?.dispose();
|
||||
lines.close();
|
||||
}
|
||||
1
src/muse/基础设施/宿主/__init__.py
Normal file
1
src/muse/基础设施/宿主/__init__.py
Normal file
@ -0,0 +1 @@
|
||||
"""按明确协议装配宿主,无导入副作用。"""
|
||||
148
src/muse/基础设施/宿主/直接调用.py
Normal file
148
src/muse/基础设施/宿主/直接调用.py
Normal file
@ -0,0 +1,148 @@
|
||||
"""直接 HTTP 的请求封装;模型与输出是否可用由 S02 后验判断。"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
from collections.abc import Awaitable, Callable
|
||||
from dataclasses import dataclass
|
||||
from typing import Any
|
||||
|
||||
from muse.任务运行.接口 import 工具请求, 模型协议错误, 模型结果, 模型请求
|
||||
from muse.基础设施.模型 import ChatCompletions, Messages, Responses
|
||||
from muse.基础设施.模型.HTTP传输 import HTTP传输, 已准备HTTP请求
|
||||
|
||||
|
||||
@dataclass(frozen=True)
|
||||
class 直接宿主:
|
||||
传输: HTTP传输
|
||||
|
||||
def 名称(self) -> str:
|
||||
return "direct"
|
||||
|
||||
def 准备(self, 请求: 模型请求) -> _直接请求:
|
||||
协议 = self.传输.协议
|
||||
for 工具 in 请求.工具:
|
||||
if not 工具.get("name") or not isinstance(工具.get("parameters"), dict):
|
||||
raise 模型协议错误("工具缺少固定名称或参数合同")
|
||||
if 请求.thinking and 协议 == "chat-completions":
|
||||
raise 模型协议错误("该直接协议的推理参数尚未登记,不能静默忽略")
|
||||
data: dict[str, Any] = {"model": 请求.model, "stream": True}
|
||||
if 协议 == "responses":
|
||||
data.update(
|
||||
{
|
||||
"instructions": 请求.系统提示,
|
||||
"input": 请求.用户输入,
|
||||
"max_output_tokens": 请求.最大输出token,
|
||||
"store": False,
|
||||
}
|
||||
)
|
||||
if 请求.thinking:
|
||||
data["reasoning"] = {"effort": 请求.thinking}
|
||||
解析 = Responses.解析响应
|
||||
elif 协议 in {"anthropic", "chat-completions"}:
|
||||
data.update(
|
||||
{
|
||||
"max_tokens": 请求.最大输出token,
|
||||
"messages": [{"role": "user", "content": 请求.用户输入}],
|
||||
}
|
||||
)
|
||||
if 协议 == "anthropic":
|
||||
data["system"] = 请求.系统提示
|
||||
if 请求.thinking == "off":
|
||||
data["thinking"] = {"type": "disabled"}
|
||||
elif 请求.thinking is not None:
|
||||
if 请求.thinking not in {"low", "medium", "high", "xhigh", "max"}:
|
||||
raise 模型协议错误("Messages 协议不支持该推理等级")
|
||||
data["thinking"] = {"type": "adaptive"}
|
||||
data["output_config"] = {"effort": 请求.thinking}
|
||||
解析 = Messages.解析响应
|
||||
else:
|
||||
data["messages"].insert(0, {"role": "system", "content": 请求.系统提示})
|
||||
data["stream_options"] = {"include_usage": True}
|
||||
解析 = ChatCompletions.解析响应
|
||||
else:
|
||||
raise 模型协议错误("未登记的直接调用协议")
|
||||
if 请求.工具:
|
||||
if 协议 == "anthropic":
|
||||
data["tools"] = [
|
||||
{
|
||||
"name": t["name"],
|
||||
"description": t.get("description", ""),
|
||||
"input_schema": t["parameters"],
|
||||
}
|
||||
for t in 请求.工具
|
||||
]
|
||||
elif 协议 == "responses":
|
||||
data["tools"] = [{"type": "function", **t} for t in 请求.工具]
|
||||
else:
|
||||
data["tools"] = [{"type": "function", "function": t} for t in 请求.工具]
|
||||
if 请求.历史:
|
||||
消息: list[dict[str, Any]] = (
|
||||
[{"role": "user", "content": 请求.用户输入}]
|
||||
if 协议 == "responses"
|
||||
else data["messages"]
|
||||
)
|
||||
for 回合 in 请求.历史:
|
||||
if 回合.模型.回放协议 != 协议 or not 回合.模型.回放内容:
|
||||
raise 模型协议错误("前序模型回合缺少当前协议的完整回放内容")
|
||||
if 协议 == "anthropic":
|
||||
消息.append({"role": "assistant", "content": list(回合.模型.回放内容)})
|
||||
if 回合.工具:
|
||||
消息.append(
|
||||
{
|
||||
"role": "user",
|
||||
"content": [
|
||||
{
|
||||
"type": "tool_result",
|
||||
"tool_use_id": t.请求.调用ID,
|
||||
"content": t.内容,
|
||||
}
|
||||
for t in 回合.工具
|
||||
],
|
||||
}
|
||||
)
|
||||
elif 协议 == "responses":
|
||||
消息.extend(回合.模型.回放内容)
|
||||
消息.extend(
|
||||
{"type": "function_call_output", "call_id": t.请求.调用ID, "output": t.内容}
|
||||
for t in 回合.工具
|
||||
)
|
||||
else:
|
||||
消息.extend(回合.模型.回放内容)
|
||||
消息.extend(
|
||||
{"role": "tool", "tool_call_id": t.请求.调用ID, "content": t.内容}
|
||||
for t in 回合.工具
|
||||
)
|
||||
data["input" if 协议 == "responses" else "messages"] = 消息
|
||||
return _直接请求(self.传输.准备(data, 请求.总期限秒), 解析)
|
||||
|
||||
async def 调用(self, 请求: 模型请求, 收到片段: Callable[[str], None] | None = None) -> 模型结果:
|
||||
return await self.准备(请求).调用(收到片段)
|
||||
|
||||
|
||||
@dataclass(frozen=True)
|
||||
class _直接请求:
|
||||
请求: 已准备HTTP请求
|
||||
解析: Callable[[list[dict]], 模型结果]
|
||||
|
||||
async def 调用(self, 收到片段: Callable[[str], None] | None = None) -> 模型结果:
|
||||
事件 = []
|
||||
async for 项 in self.请求.流():
|
||||
事件.append(项)
|
||||
if 收到片段 is not None and 项.get("type") == "response.output_text.delta":
|
||||
收到片段(项.get("delta", ""))
|
||||
return self.解析(事件)
|
||||
|
||||
|
||||
class 直接角色循环:
|
||||
async def 运行(
|
||||
self,
|
||||
请求: 模型请求,
|
||||
模型回合: Callable[[], Awaitable[模型结果]],
|
||||
工具回合: Callable[[工具请求], Awaitable[str]],
|
||||
) -> 模型结果:
|
||||
while True:
|
||||
结果 = await 模型回合()
|
||||
if 结果.状态 != "tool_calls":
|
||||
return 结果
|
||||
for 工具 in 结果.工具调用:
|
||||
await 工具回合(工具)
|
||||
74
src/muse/基础设施/模型/ChatCompletions.py
Normal file
74
src/muse/基础设施/模型/ChatCompletions.py
Normal file
@ -0,0 +1,74 @@
|
||||
"""Chat Completions 必须有 finish_reason,流结束标记自身不证明完成。"""
|
||||
|
||||
import json
|
||||
from collections.abc import Iterable
|
||||
from typing import Any
|
||||
|
||||
from muse.任务运行.接口 import 工具请求, 模型协议错误, 模型用量, 模型结果
|
||||
|
||||
|
||||
def 解析响应(事件: Iterable[dict]) -> 模型结果:
|
||||
模型 = 身份 = 原因 = None
|
||||
文本, 用量, 工具 = [], None, {}
|
||||
for 项 in 事件:
|
||||
if 项.get("error"):
|
||||
return 模型结果("failed", "".join(文本), 模型, 用量, 身份, 失败码="PROVIDER_ERROR")
|
||||
模型 = 项.get("model", 模型)
|
||||
身份 = 项.get("id", 身份)
|
||||
if 项.get("usage"):
|
||||
u = 项["usage"]
|
||||
用量 = 模型用量(
|
||||
u["prompt_tokens"],
|
||||
u["completion_tokens"],
|
||||
(u.get("prompt_tokens_details") or {}).get("cached_tokens", 0),
|
||||
)
|
||||
for choice in 项.get("choices", []):
|
||||
if choice.get("index", 0) != 0:
|
||||
raise 模型协议错误("一次角色调用只接受一个候选分支")
|
||||
原因 = choice.get("finish_reason") or 原因
|
||||
delta = choice.get("delta", {})
|
||||
if delta.get("content"):
|
||||
文本.append(delta["content"])
|
||||
for call in delta.get("tool_calls", []):
|
||||
当前 = 工具.setdefault(call["index"], {"id": "", "name": "", "arguments": ""})
|
||||
当前["id"] = call.get("id", 当前["id"])
|
||||
for key in ("name", "arguments"):
|
||||
当前[key] += call.get("function", {}).get(key, "")
|
||||
调用 = []
|
||||
for _, call in sorted(工具.items()):
|
||||
try:
|
||||
参数 = json.loads(call["arguments"])
|
||||
if not isinstance(参数, dict):
|
||||
raise ValueError()
|
||||
调用.append(工具请求(call["id"], call["name"], 参数))
|
||||
except ValueError:
|
||||
raise 模型协议错误("Chat Completions 工具参数无效") from None
|
||||
状态 = (
|
||||
"completed"
|
||||
if 原因 == "stop"
|
||||
else "tool_calls"
|
||||
if 原因 == "tool_calls" and 调用
|
||||
else "incomplete"
|
||||
)
|
||||
if 原因 == "content_filter":
|
||||
状态 = "failed"
|
||||
消息: dict[str, Any] = {"role": "assistant", "content": "".join(文本)}
|
||||
if 调用:
|
||||
消息["tool_calls"] = [
|
||||
{
|
||||
"id": t.调用ID,
|
||||
"type": "function",
|
||||
"function": {"name": t.名称, "arguments": json.dumps(t.参数, ensure_ascii=False)},
|
||||
}
|
||||
for t in 调用
|
||||
]
|
||||
return 模型结果(
|
||||
状态,
|
||||
"".join(文本),
|
||||
模型,
|
||||
用量,
|
||||
身份,
|
||||
tuple(调用),
|
||||
回放协议="chat-completions",
|
||||
回放内容=(消息,),
|
||||
)
|
||||
104
src/muse/基础设施/模型/HTTP传输.py
Normal file
104
src/muse/基础设施/模型/HTTP传输.py
Normal file
@ -0,0 +1,104 @@
|
||||
"""异步 HTTP 与按字节分隔的 SSE;整个请求共享总期限且不隐式重试。"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import asyncio
|
||||
import json
|
||||
from collections.abc import AsyncIterable, AsyncIterator
|
||||
from dataclasses import dataclass, field
|
||||
|
||||
import httpx
|
||||
|
||||
from muse.任务运行.接口 import 模型协议错误
|
||||
from muse.基础设施.凭据读取 import 读取凭据
|
||||
|
||||
|
||||
async def 解码SSE(字节流: AsyncIterable[bytes]) -> AsyncIterator[dict]:
|
||||
缓冲 = b""
|
||||
数据: list[bytes] = []
|
||||
async for 块 in 字节流:
|
||||
缓冲 += 块
|
||||
while b"\n" in 缓冲:
|
||||
行, 缓冲 = 缓冲.split(b"\n", 1)
|
||||
行 = 行.removesuffix(b"\r")
|
||||
if 行.startswith(b"data:"):
|
||||
数据.append(行[5:].removeprefix(b" "))
|
||||
elif not 行 and 数据:
|
||||
文本 = b"\n".join(数据)
|
||||
数据.clear()
|
||||
if 文本 == b"[DONE]":
|
||||
continue
|
||||
try:
|
||||
对象 = json.loads(文本.decode("utf-8"))
|
||||
except (UnicodeError, ValueError):
|
||||
raise 模型协议错误("SSE 帧不是合法 UTF-8 JSON") from None
|
||||
if not isinstance(对象, dict):
|
||||
raise 模型协议错误("SSE 事件必须是对象")
|
||||
yield 对象
|
||||
if 缓冲.strip() or 数据:
|
||||
raise 模型协议错误("SSE 在完整事件分隔符之前断开")
|
||||
|
||||
|
||||
@dataclass(frozen=True)
|
||||
class HTTP传输:
|
||||
地址: str
|
||||
凭据方式: str
|
||||
凭据位置: str
|
||||
协议: str
|
||||
|
||||
def 准备(self, 请求: dict, 总期限秒: float) -> 已准备HTTP请求:
|
||||
try:
|
||||
地址 = httpx.URL(self.地址)
|
||||
if (
|
||||
地址.scheme not in {"http", "https"}
|
||||
or not 地址.host
|
||||
or 地址.userinfo
|
||||
or 地址.fragment
|
||||
):
|
||||
raise ValueError
|
||||
数据 = json.dumps(请求, ensure_ascii=False, allow_nan=False).encode("utf-8")
|
||||
except (ValueError, TypeError, httpx.InvalidURL):
|
||||
raise 模型协议错误("模型地址或请求序列化无效") from None
|
||||
if self.协议 not in {"responses", "anthropic", "chat-completions"}:
|
||||
raise 模型协议错误("未登记的模型传输协议")
|
||||
if 总期限秒 <= 0:
|
||||
raise 模型协议错误("模型调用期限必须为正数")
|
||||
密钥 = 读取凭据(self.凭据方式, self.凭据位置)
|
||||
头 = {"Content-Type": "application/json"}
|
||||
if self.协议 == "anthropic":
|
||||
头.update({"x-api-key": 密钥, "anthropic-version": "2023-06-01"})
|
||||
else:
|
||||
头["Authorization"] = "Bearer " + 密钥
|
||||
return 已准备HTTP请求(str(地址), tuple(头.items()), 数据, 总期限秒)
|
||||
|
||||
async def 流(self, 请求: dict, 总期限秒: float) -> AsyncIterator[dict]:
|
||||
async for 事件 in self.准备(请求, 总期限秒).流():
|
||||
yield 事件
|
||||
|
||||
|
||||
@dataclass(frozen=True)
|
||||
class 已准备HTTP请求:
|
||||
地址: str
|
||||
头: tuple[tuple[str, str], ...] = field(repr=False)
|
||||
数据: bytes = field(repr=False)
|
||||
总期限秒: float
|
||||
|
||||
async def 流(self) -> AsyncIterator[dict]:
|
||||
try:
|
||||
async with asyncio.timeout(self.总期限秒):
|
||||
async with httpx.AsyncClient(
|
||||
timeout=self.总期限秒, follow_redirects=False
|
||||
) as 客户端:
|
||||
async with 客户端.stream(
|
||||
"POST", self.地址, content=self.数据, headers=dict(self.头)
|
||||
) as 响应:
|
||||
if 响应.is_error:
|
||||
raise 模型协议错误(
|
||||
"模型 HTTP 请求失败", 上下文={"status": 响应.status_code}
|
||||
)
|
||||
async for 事件 in 解码SSE(响应.aiter_bytes()):
|
||||
yield 事件
|
||||
except (TimeoutError, httpx.TimeoutException):
|
||||
raise 模型协议错误("模型调用超过总期限") from None
|
||||
except httpx.HTTPError:
|
||||
raise 模型协议错误("模型传输中断,调用结果可能未知") from None
|
||||
77
src/muse/基础设施/模型/Messages.py
Normal file
77
src/muse/基础设施/模型/Messages.py
Normal file
@ -0,0 +1,77 @@
|
||||
"""Anthropic message_start/message_stop 与 stop_reason 一起确定终态。"""
|
||||
|
||||
import json
|
||||
from collections.abc import Iterable
|
||||
|
||||
from muse.任务运行.接口 import 工具请求, 模型协议错误, 模型用量, 模型结果
|
||||
|
||||
|
||||
def 解析响应(事件: Iterable[dict]) -> 模型结果:
|
||||
消息, 块, 用量 = {}, {}, {}
|
||||
开始 = 结束 = False
|
||||
原因 = None
|
||||
for 项 in 事件:
|
||||
类型 = 项.get("type")
|
||||
if 类型 == "message_start":
|
||||
消息 = 项["message"]
|
||||
用量.update(消息.get("usage") or {})
|
||||
开始 = True
|
||||
elif 类型 == "content_block_start":
|
||||
块[项["index"]] = dict(项["content_block"])
|
||||
elif 类型 == "content_block_delta":
|
||||
目标 = 块[项["index"]]
|
||||
delta = 项["delta"]
|
||||
if delta["type"] == "text_delta":
|
||||
目标["text"] = 目标.get("text", "") + delta["text"]
|
||||
elif delta["type"] == "input_json_delta":
|
||||
目标["partial_json"] = 目标.get("partial_json", "") + delta["partial_json"]
|
||||
elif delta["type"] == "thinking_delta":
|
||||
目标["thinking"] = 目标.get("thinking", "") + delta["thinking"]
|
||||
elif delta["type"] == "signature_delta":
|
||||
目标["signature"] = 目标.get("signature", "") + delta["signature"]
|
||||
elif 类型 == "message_delta":
|
||||
原因 = 项.get("delta", {}).get("stop_reason", 原因)
|
||||
用量.update(项.get("usage") or {})
|
||||
elif 类型 == "message_stop":
|
||||
结束 = True
|
||||
elif 类型 == "error":
|
||||
return 模型结果("failed", "", 消息.get("model"), None, 失败码="PROVIDER_ERROR")
|
||||
文本 = "".join(b.get("text", "") for _, b in sorted(块.items()) if b.get("type") == "text")
|
||||
工具 = []
|
||||
for b in 块.values():
|
||||
if b.get("type") == "tool_use":
|
||||
try:
|
||||
参数 = json.loads(b["partial_json"]) if "partial_json" in b else b["input"]
|
||||
if not isinstance(参数, dict):
|
||||
raise ValueError()
|
||||
工具.append(工具请求(b["id"], b["name"], 参数))
|
||||
b["input"] = 参数
|
||||
b.pop("partial_json", None)
|
||||
except (ValueError, KeyError):
|
||||
raise 模型协议错误("Anthropic 工具参数无效") from None
|
||||
统计 = None
|
||||
if "input_tokens" in 用量 and "output_tokens" in 用量:
|
||||
统计 = 模型用量(
|
||||
用量["input_tokens"],
|
||||
用量["output_tokens"],
|
||||
用量.get("cache_read_input_tokens", 0),
|
||||
用量.get("cache_creation_input_tokens", 0),
|
||||
)
|
||||
状态 = "incomplete"
|
||||
if 开始 and 结束:
|
||||
if 原因 in {"end_turn", "stop_sequence"}:
|
||||
状态 = "completed"
|
||||
elif 原因 == "tool_use" and 工具:
|
||||
状态 = "tool_calls"
|
||||
elif 原因 in {"refusal", "error"}:
|
||||
状态 = "failed"
|
||||
return 模型结果(
|
||||
状态,
|
||||
文本,
|
||||
消息.get("model"),
|
||||
统计,
|
||||
消息.get("id"),
|
||||
tuple(工具),
|
||||
回放协议="anthropic",
|
||||
回放内容=tuple(b for _, b in sorted(块.items())),
|
||||
)
|
||||
65
src/muse/基础设施/模型/Responses.py
Normal file
65
src/muse/基础设施/模型/Responses.py
Normal file
@ -0,0 +1,65 @@
|
||||
"""Responses 的完成对象是最终依据;delta 仅用于过程显示。"""
|
||||
|
||||
import json
|
||||
from collections.abc import Iterable
|
||||
|
||||
from muse.任务运行.接口 import 工具请求, 模型协议错误, 模型用量, 模型结果
|
||||
|
||||
|
||||
def 解析响应(事件: Iterable[dict]) -> 模型结果:
|
||||
响应 = None
|
||||
片段 = []
|
||||
for 项 in 事件:
|
||||
类型 = 项.get("type")
|
||||
if 类型 == "response.output_text.delta":
|
||||
片段.append(项.get("delta", ""))
|
||||
elif 类型 in {"response.completed", "response.failed", "response.incomplete"}:
|
||||
响应 = 项.get("response")
|
||||
if not isinstance(响应, dict):
|
||||
raise 模型协议错误("Responses 终态缺少响应对象")
|
||||
if 响应.get("status") != 类型.split(".")[1]:
|
||||
raise 模型协议错误("Responses 事件与响应状态不一致")
|
||||
elif 类型 == "error":
|
||||
return 模型结果("failed", "".join(片段), None, None, 失败码=项.get("code"))
|
||||
if 响应 is None:
|
||||
return 模型结果("incomplete", "".join(片段), None, None, 失败码="TERMINAL_MISSING")
|
||||
状态 = 响应["status"]
|
||||
文本, 工具 = [], []
|
||||
for 输出 in 响应.get("output", []):
|
||||
if 输出.get("type") == "message":
|
||||
for 内容 in 输出.get("content", []):
|
||||
if 内容.get("type") == "output_text":
|
||||
文本.append(内容["text"])
|
||||
elif 内容.get("type") == "refusal":
|
||||
状态 = "failed"
|
||||
elif 输出.get("type") == "function_call":
|
||||
try:
|
||||
参数 = json.loads(输出["arguments"])
|
||||
if not isinstance(参数, dict):
|
||||
raise ValueError()
|
||||
工具.append(工具请求(输出["call_id"], 输出["name"], 参数))
|
||||
except (ValueError, KeyError):
|
||||
raise 模型协议错误("模型工具调用参数不符合对象合同") from None
|
||||
用量 = 响应.get("usage")
|
||||
统计 = None
|
||||
if 用量 is not None:
|
||||
try:
|
||||
统计 = 模型用量(
|
||||
用量["input_tokens"],
|
||||
用量["output_tokens"],
|
||||
(用量.get("input_tokens_details") or {}).get("cached_tokens", 0),
|
||||
)
|
||||
except (KeyError, TypeError):
|
||||
raise 模型协议错误("Responses 用量字段不完整") from None
|
||||
if 工具 and 状态 == "completed":
|
||||
状态 = "tool_calls"
|
||||
return 模型结果(
|
||||
状态,
|
||||
"".join(文本),
|
||||
响应.get("model"),
|
||||
统计,
|
||||
响应.get("id"),
|
||||
tuple(工具),
|
||||
回放协议="responses",
|
||||
回放内容=tuple(响应.get("output", [])),
|
||||
)
|
||||
1
src/muse/基础设施/模型/__init__.py
Normal file
1
src/muse/基础设施/模型/__init__.py
Normal file
@ -0,0 +1 @@
|
||||
"""协议适配,不在导入时建立网络连接。"""
|
||||
42
src/muse/接入/cli/管理命令.py
Normal file
42
src/muse/接入/cli/管理命令.py
Normal file
@ -0,0 +1,42 @@
|
||||
"""管理具名运行配置草案;不将保存或离线检查当作生产验证启用。"""
|
||||
|
||||
import tomllib
|
||||
from dataclasses import asdict
|
||||
from pathlib import Path
|
||||
|
||||
from muse.任务运行.接口 import 运行配置内容, 配置版本管理
|
||||
from muse.共享.错误 import 配置错误
|
||||
from muse.启动 import 应用装配
|
||||
from muse.基础设施.数据库.连接 import 数据库工厂
|
||||
|
||||
|
||||
def 运行配置命令(
|
||||
装配: 应用装配,
|
||||
动作: str,
|
||||
配置ID: str,
|
||||
版本: str,
|
||||
内容文件: str | None,
|
||||
) -> dict:
|
||||
管理 = 配置版本管理(数据库工厂(装配.配置.数据库, 装配.配置.运行用途))
|
||||
if 动作 == "保存":
|
||||
if 内容文件 is None:
|
||||
raise 配置错误("保存配置草案需要显式提供TOML内容文件")
|
||||
try:
|
||||
内容 = 运行配置内容.从快照(tomllib.loads(Path(内容文件).read_text()))
|
||||
except (OSError, ValueError, TypeError, KeyError):
|
||||
raise 配置错误("运行配置草案不是完整的已定义TOML结构") from None
|
||||
快照 = 管理.保存草案(配置ID, 版本, 内容)
|
||||
elif 动作 == "查看":
|
||||
快照 = 管理.读取版本(配置ID, 版本)
|
||||
else:
|
||||
raise 配置错误("未支持的运行配置管理动作")
|
||||
return asdict(快照)
|
||||
|
||||
|
||||
def 运行管理命令(装配: 应用装配, 动作: str, 对象ID: str | None) -> dict:
|
||||
原文 = 装配.要求原文()
|
||||
if 动作 == "清理无租约暂存":
|
||||
return 原文.清理无租约原文()
|
||||
if 动作 == "查看清理回执" and 对象ID:
|
||||
return 原文.读取孤儿清理回执(对象ID)
|
||||
raise 配置错误("管理动作或对应对象身份不完整")
|
||||
85
tests/单元/test_预算与配置.py
Normal file
85
tests/单元/test_预算与配置.py
Normal file
@ -0,0 +1,85 @@
|
||||
"""预算时间与金额、配置引用的确定性合同。"""
|
||||
|
||||
from datetime import datetime, timedelta
|
||||
from decimal import Decimal
|
||||
from zoneinfo import ZoneInfo
|
||||
|
||||
import pytest
|
||||
|
||||
from muse.任务运行.配置版本 import 凭据引用, 提供方配置, 运行配置内容, 配置版本错误
|
||||
from muse.任务运行.预算管理 import 任务预算计划, 归属窗口, 角色预算, 预算错误, 额度策略
|
||||
|
||||
|
||||
@pytest.mark.parametrize(
|
||||
"hour,expected",
|
||||
[(0, 0), (4, 0), (5, 5), (9, 5), (10, 10), (14, 10), (15, 15), (19, 15), (20, 20), (23, 20)],
|
||||
ids=["00", "04", "05", "09", "10", "14", "15", "19", "20", "23"],
|
||||
)
|
||||
def test_五小时额度窗口承接每日边界__2852b5(hour: int, expected: int) -> None:
|
||||
当前 = datetime(2026, 7, 16, hour, 30, tzinfo=ZoneInfo("Asia/Shanghai"))
|
||||
窗口 = 归属窗口(当前)
|
||||
assert 窗口.起点.strftime("%Y-%m-%dT%H") == f"2026-07-16T{expected:02}"
|
||||
assert 窗口.终点 - 窗口.起点 == timedelta(hours=4 if hour >= 20 else 5)
|
||||
|
||||
|
||||
@pytest.mark.parametrize(
|
||||
"hour,minute,seconds",
|
||||
[(3, 0, 7200), (19, 30, 1800), (21, 0, 10800), (23, 59, 60)],
|
||||
ids=["morning", "20-boundary", "midnight", "last-minute"],
|
||||
)
|
||||
def test_下一窗口等待秒数按原日界线计算__7790f8(hour: int, minute: int, seconds: int) -> None:
|
||||
当前 = datetime(2026, 9, 9, hour, minute, tzinfo=ZoneInfo("Asia/Shanghai"))
|
||||
assert (归属窗口(当前).终点 - 当前).total_seconds() == seconds
|
||||
|
||||
|
||||
def test_预算计划保留调用计划并按单次上限预留__a60003() -> None:
|
||||
角色 = (
|
||||
角色预算("writer", 45, 150, Decimal("5")),
|
||||
角色预算("detector", 24, 150, Decimal("5")),
|
||||
角色预算("judge", 45, 150, Decimal("5")),
|
||||
)
|
||||
计划 = 任务预算计划(
|
||||
Decimal("570"), 角色, "approval-1", datetime(2030, 1, 1, tzinfo=ZoneInfo("UTC"))
|
||||
)
|
||||
assert 计划.最坏预留 == Decimal("570.000000")
|
||||
assert 计划.冻结()["角色"][0]["计划次数"] == 45
|
||||
with pytest.raises(预算错误):
|
||||
任务预算计划(Decimal("569.99"), 角色, "approval-1", 计划.截止时间)
|
||||
|
||||
|
||||
@pytest.mark.parametrize(
|
||||
"amount",
|
||||
[Decimal("NaN"), Decimal("-1"), Decimal("Infinity")],
|
||||
ids=["nan", "negative", "infinite"],
|
||||
)
|
||||
def test_预算拒绝不确定或负数金额__a60001(amount: Decimal) -> None:
|
||||
with pytest.raises(预算错误):
|
||||
额度策略("quota", "1", amount, 6000)
|
||||
|
||||
|
||||
def test_配置凭据只保留引用且调用参数不可藏入明文__a60002() -> None:
|
||||
引用 = 凭据引用("api", "环境变量", "MUSE_PROVIDER_KEY")
|
||||
内容 = 运行配置内容(
|
||||
"direct",
|
||||
"1",
|
||||
"role-policy-1",
|
||||
"release-1",
|
||||
"quota-1",
|
||||
{"writer": {"provider": "provider", "model": "model", "thinking": "high"}},
|
||||
(引用,),
|
||||
(提供方配置("provider", "responses", "https://provider.invalid/v1/responses", "api"),),
|
||||
"price-1",
|
||||
)
|
||||
assert 内容.冻结()["凭据"][0]["位置"] == "MUSE_PROVIDER_KEY"
|
||||
with pytest.raises(配置版本错误):
|
||||
运行配置内容(
|
||||
"direct",
|
||||
"1",
|
||||
"policy",
|
||||
"release",
|
||||
"quota",
|
||||
{"writer": {"api_key": "synthetic-secret"}},
|
||||
(),
|
||||
(),
|
||||
"price-1",
|
||||
)
|
||||
205
tests/夹具/Pi合成事件.json
Normal file
205
tests/夹具/Pi合成事件.json
Normal file
@ -0,0 +1,205 @@
|
||||
[
|
||||
{
|
||||
"type": "agent_start"
|
||||
},
|
||||
{
|
||||
"type": "message_end",
|
||||
"message": {
|
||||
"role": "user",
|
||||
"content": [
|
||||
{
|
||||
"type": "text",
|
||||
"text": "读结构后返回JSON。"
|
||||
}
|
||||
],
|
||||
"timestamp": 1788965416161
|
||||
}
|
||||
},
|
||||
{
|
||||
"type": "message_end",
|
||||
"message": {
|
||||
"role": "assistant",
|
||||
"content": [
|
||||
{
|
||||
"type": "toolCall",
|
||||
"id": "probe-call",
|
||||
"name": "schema_read",
|
||||
"arguments": {}
|
||||
}
|
||||
],
|
||||
"api": "anthropic-messages",
|
||||
"provider": "muse-probe",
|
||||
"model": "synthetic-model",
|
||||
"usage": {
|
||||
"input": 1,
|
||||
"output": 1,
|
||||
"cacheRead": 0,
|
||||
"cacheWrite": 0,
|
||||
"totalTokens": 2,
|
||||
"cost": {
|
||||
"input": 0,
|
||||
"output": 0,
|
||||
"cacheRead": 0,
|
||||
"cacheWrite": 0,
|
||||
"total": 0
|
||||
}
|
||||
},
|
||||
"stopReason": "toolUse",
|
||||
"timestamp": 1788965416165
|
||||
}
|
||||
},
|
||||
{
|
||||
"type": "tool_execution_start",
|
||||
"toolCallId": "probe-call",
|
||||
"toolName": "schema_read"
|
||||
},
|
||||
{
|
||||
"type": "tool_execution_end",
|
||||
"toolCallId": "probe-call",
|
||||
"toolName": "schema_read",
|
||||
"isError": false
|
||||
},
|
||||
{
|
||||
"type": "message_end",
|
||||
"message": {
|
||||
"role": "toolResult",
|
||||
"toolCallId": "probe-call",
|
||||
"toolName": "schema_read",
|
||||
"content": [
|
||||
{
|
||||
"type": "text",
|
||||
"text": "{\"type\":\"object\"}"
|
||||
}
|
||||
],
|
||||
"details": {
|
||||
"synthetic": true
|
||||
},
|
||||
"isError": false,
|
||||
"timestamp": 1788965416167
|
||||
}
|
||||
},
|
||||
{
|
||||
"type": "message_end",
|
||||
"message": {
|
||||
"role": "assistant",
|
||||
"content": [
|
||||
{
|
||||
"type": "text",
|
||||
"text": "{\"text\":\"合成工具已返回\"}"
|
||||
}
|
||||
],
|
||||
"api": "anthropic-messages",
|
||||
"provider": "muse-probe",
|
||||
"model": "synthetic-model",
|
||||
"usage": {
|
||||
"input": 1,
|
||||
"output": 1,
|
||||
"cacheRead": 0,
|
||||
"cacheWrite": 0,
|
||||
"totalTokens": 2,
|
||||
"cost": {
|
||||
"input": 0,
|
||||
"output": 0,
|
||||
"cacheRead": 0,
|
||||
"cacheWrite": 0,
|
||||
"total": 0
|
||||
}
|
||||
},
|
||||
"stopReason": "stop",
|
||||
"timestamp": 1788965416167
|
||||
}
|
||||
},
|
||||
{
|
||||
"type": "agent_end",
|
||||
"messages": [
|
||||
{
|
||||
"role": "user",
|
||||
"content": [
|
||||
{
|
||||
"type": "text",
|
||||
"text": "读结构后返回JSON。"
|
||||
}
|
||||
],
|
||||
"timestamp": 1788965416161
|
||||
},
|
||||
{
|
||||
"role": "assistant",
|
||||
"content": [
|
||||
{
|
||||
"type": "toolCall",
|
||||
"id": "probe-call",
|
||||
"name": "schema_read",
|
||||
"arguments": {}
|
||||
}
|
||||
],
|
||||
"api": "anthropic-messages",
|
||||
"provider": "muse-probe",
|
||||
"model": "synthetic-model",
|
||||
"usage": {
|
||||
"input": 1,
|
||||
"output": 1,
|
||||
"cacheRead": 0,
|
||||
"cacheWrite": 0,
|
||||
"totalTokens": 2,
|
||||
"cost": {
|
||||
"input": 0,
|
||||
"output": 0,
|
||||
"cacheRead": 0,
|
||||
"cacheWrite": 0,
|
||||
"total": 0
|
||||
}
|
||||
},
|
||||
"stopReason": "toolUse",
|
||||
"timestamp": 1788965416165
|
||||
},
|
||||
{
|
||||
"role": "toolResult",
|
||||
"toolCallId": "probe-call",
|
||||
"toolName": "schema_read",
|
||||
"content": [
|
||||
{
|
||||
"type": "text",
|
||||
"text": "{\"type\":\"object\"}"
|
||||
}
|
||||
],
|
||||
"details": {
|
||||
"synthetic": true
|
||||
},
|
||||
"isError": false,
|
||||
"timestamp": 1788965416167
|
||||
},
|
||||
{
|
||||
"role": "assistant",
|
||||
"content": [
|
||||
{
|
||||
"type": "text",
|
||||
"text": "{\"text\":\"合成工具已返回\"}"
|
||||
}
|
||||
],
|
||||
"api": "anthropic-messages",
|
||||
"provider": "muse-probe",
|
||||
"model": "synthetic-model",
|
||||
"usage": {
|
||||
"input": 1,
|
||||
"output": 1,
|
||||
"cacheRead": 0,
|
||||
"cacheWrite": 0,
|
||||
"totalTokens": 2,
|
||||
"cost": {
|
||||
"input": 0,
|
||||
"output": 0,
|
||||
"cacheRead": 0,
|
||||
"cacheWrite": 0,
|
||||
"total": 0
|
||||
}
|
||||
},
|
||||
"stopReason": "stop",
|
||||
"timestamp": 1788965416167
|
||||
}
|
||||
],
|
||||
"willRetry": false
|
||||
},
|
||||
{
|
||||
"type": "agent_settled"
|
||||
}
|
||||
]
|
||||
119
tests/契约/test_宿主执行合同.py
Normal file
119
tests/契约/test_宿主执行合同.py
Normal file
@ -0,0 +1,119 @@
|
||||
"""实际宿主封装配合显式传输替身;不作为真实模型调用证据。"""
|
||||
|
||||
import asyncio
|
||||
|
||||
import pytest
|
||||
|
||||
from muse.任务运行.接口 import 模型协议错误, 模型请求
|
||||
from muse.基础设施.宿主.直接调用 import 直接宿主
|
||||
|
||||
|
||||
class 合成传输:
|
||||
协议 = "anthropic"
|
||||
|
||||
def __init__(self):
|
||||
self.收到 = []
|
||||
|
||||
def 准备(self, 请求, 总期限秒):
|
||||
self.请求, self.期限 = 请求, 总期限秒
|
||||
return self
|
||||
|
||||
async def 流(self):
|
||||
请求, 总期限秒 = self.请求, self.期限
|
||||
self.收到.append((请求, 总期限秒))
|
||||
yield {
|
||||
"type": "message_start",
|
||||
"message": {"id": "m1", "model": "claude-opus-4-8", "usage": {"input_tokens": 1}},
|
||||
}
|
||||
yield {
|
||||
"type": "content_block_start",
|
||||
"index": 0,
|
||||
"content_block": {"type": "text", "text": "{}"},
|
||||
}
|
||||
yield {
|
||||
"type": "message_delta",
|
||||
"delta": {"stop_reason": "end_turn"},
|
||||
"usage": {"output_tokens": 1},
|
||||
}
|
||||
yield {"type": "message_stop"}
|
||||
|
||||
|
||||
def test_直接宿主传递显式模型期限与推理等级__a64001() -> None:
|
||||
传输 = 合成传输()
|
||||
请求 = 模型请求(
|
||||
"call-1",
|
||||
"configured",
|
||||
"claude-opus-4-8[1M]",
|
||||
"系统说明",
|
||||
"合成输入",
|
||||
{"type": "object"},
|
||||
1024,
|
||||
30,
|
||||
thinking="high",
|
||||
)
|
||||
结果 = asyncio.run(直接宿主(传输).调用(请求))
|
||||
已发, 秒 = 传输.收到[0]
|
||||
assert 已发["model"] == 请求.model and 秒 == 30
|
||||
assert 已发["thinking"] == {"type": "adaptive"}
|
||||
assert 已发["output_config"] == {"effort": "high"}
|
||||
assert 结果.状态 == "completed" and 结果.实际模型 == "claude-opus-4-8"
|
||||
|
||||
|
||||
def test_无工具直接宿主拒绝请求工具而不静默降级__a64002() -> None:
|
||||
传输 = 合成传输()
|
||||
请求 = 模型请求(
|
||||
"call-1",
|
||||
"configured",
|
||||
"claude-opus-4-8[1M]",
|
||||
"系统说明",
|
||||
"合成输入",
|
||||
{"type": "object"},
|
||||
1024,
|
||||
30,
|
||||
工具=({"name": "read"},),
|
||||
)
|
||||
with pytest.raises(模型协议错误):
|
||||
asyncio.run(直接宿主(传输).调用(请求))
|
||||
assert not 传输.收到
|
||||
|
||||
|
||||
def test_Pi真实事件保留工具身份并以已结算回合收尾__a64003():
|
||||
import json
|
||||
from pathlib import Path
|
||||
|
||||
from muse.任务运行.接口 import 模型用量, 模型结果
|
||||
from muse.基础设施.宿主.Pi事件归一 import 归一事件, 检查会话结束
|
||||
|
||||
事件 = json.loads((Path(__file__).parents[1] / "夹具/Pi合成事件.json").read_text())
|
||||
工具 = [v for e in 事件 for v in 归一事件(e) if v["type"] == "tool.completed"]
|
||||
assert 工具 == [
|
||||
{"type": "tool.completed", "call_id": "probe-call", "tool": "schema_read", "failed": False}
|
||||
]
|
||||
回合 = 模型结果(
|
||||
"completed", '{"text":"合成工具已返回"}', "verified-upstream-model", 模型用量(1, 1)
|
||||
)
|
||||
assert 检查会话结束(事件, 0, 回合) is 回合
|
||||
with pytest.raises(模型协议错误, match="结束"):
|
||||
检查会话结束(事件[:-1], 0, 回合)
|
||||
|
||||
|
||||
def test_Pi自动续跑或改写结果不能冒充完成__a64004():
|
||||
from muse.任务运行.接口 import 模型结果
|
||||
from muse.基础设施.宿主.Pi事件归一 import 检查会话结束
|
||||
|
||||
回合 = 模型结果("completed", "原始结果", "verified-upstream-model", None)
|
||||
with pytest.raises(模型协议错误, match="未登记"):
|
||||
检查会话结束([{"type": "auto_retry_start"}], 0, 回合)
|
||||
事件 = [
|
||||
{
|
||||
"type": "message_end",
|
||||
"message": {
|
||||
"role": "assistant",
|
||||
"stopReason": "stop",
|
||||
"content": [{"type": "text", "text": "另一个结果"}],
|
||||
},
|
||||
},
|
||||
{"type": "agent_settled"},
|
||||
]
|
||||
with pytest.raises(模型协议错误, match="回合"):
|
||||
检查会话结束(事件, 0, 回合)
|
||||
64
tests/契约/test_工具范围.py
Normal file
64
tests/契约/test_工具范围.py
Normal file
@ -0,0 +1,64 @@
|
||||
"""工具范围的独立协议测试;合成读取器不代表B09真实数据库权限。"""
|
||||
|
||||
import pytest
|
||||
|
||||
from muse.任务运行.接口 import (
|
||||
只读工具集,
|
||||
工具定义,
|
||||
工具来源,
|
||||
工具结果,
|
||||
工具范围,
|
||||
工具请求,
|
||||
模型协议错误,
|
||||
)
|
||||
|
||||
|
||||
def test_工具只能消费冻结范围且返回可追溯版本__a65001() -> None:
|
||||
记录 = []
|
||||
|
||||
def 读取(范围, 条件):
|
||||
assert 范围.作品ID == "work-a" and 范围.截止位置 == 3
|
||||
assert 条件 == {"query": "桥"}
|
||||
return 工具结果("第三章的桥", (工具来源("source-a", "3", "schema1", "projection1"),))
|
||||
|
||||
定义 = 工具定义(
|
||||
"read",
|
||||
"读取指定来源",
|
||||
{
|
||||
"type": "object",
|
||||
"properties": {"query": {"type": "string"}},
|
||||
"required": ["query"],
|
||||
"additionalProperties": False,
|
||||
},
|
||||
读取,
|
||||
)
|
||||
集合 = 只读工具集(
|
||||
(定义,),
|
||||
("read",),
|
||||
工具范围("task", "work-a", ("source-a",), 3, "production", "generation"),
|
||||
记录.append,
|
||||
)
|
||||
assert 集合.调用(工具请求("call1", "read", {"query": "桥"})).内容 == "第三章的桥"
|
||||
assert 记录[0]["source_refs"][0]["revision"] == "3"
|
||||
with pytest.raises(模型协议错误):
|
||||
集合.调用(工具请求("call2", "read", {"query": "桥", "work_id": "work-b"}))
|
||||
with pytest.raises(模型协议错误):
|
||||
集合.调用(工具请求("call3", "write", {"query": "桥"}))
|
||||
assert len(记录) == 1
|
||||
|
||||
|
||||
def test_所需工具未加载和越界来源均拒绝__a65002() -> None:
|
||||
范围 = 工具范围("task", "work-a", ("source-a",), 3, "production", "generation")
|
||||
with pytest.raises(模型协议错误):
|
||||
只读工具集((), ("read",), 范围, lambda _: None)
|
||||
工具 = 工具定义(
|
||||
"read",
|
||||
"读取",
|
||||
{"type": "object"},
|
||||
lambda *_: 工具结果("不应送入模型", (工具来源("source-b", "1", "s", "p"),)),
|
||||
)
|
||||
记录 = []
|
||||
集合 = 只读工具集((工具,), ("read",), 范围, 记录.append)
|
||||
with pytest.raises(模型协议错误):
|
||||
集合.调用(工具请求("call", "read", {}))
|
||||
assert not 记录
|
||||
181
tests/契约/test_流式终态.py
Normal file
181
tests/契约/test_流式终态.py
Normal file
@ -0,0 +1,181 @@
|
||||
"""合成流验证模型协议;不发送网络请求,不作为真实模型证据。"""
|
||||
|
||||
import asyncio
|
||||
import json
|
||||
|
||||
import pytest
|
||||
|
||||
from muse.任务运行.接口 import 模型协议错误
|
||||
from muse.基础设施.模型.HTTP传输 import 解码SSE
|
||||
from muse.基础设施.模型.Responses import 解析响应
|
||||
|
||||
|
||||
def 完成事件(*, status="completed", usage=True):
|
||||
return {
|
||||
"type": "response." + status,
|
||||
"response": {
|
||||
"id": "response-1",
|
||||
"status": status,
|
||||
"model": "test-model",
|
||||
"output": [
|
||||
{
|
||||
"type": "message",
|
||||
"content": [{"type": "output_text", "text": "中文\u0085保持。"}],
|
||||
}
|
||||
],
|
||||
"usage": {"input_tokens": 10, "output_tokens": 4} if usage else None,
|
||||
},
|
||||
}
|
||||
|
||||
|
||||
def test_Responses明确终态与UTF8字节边界__a61001() -> None:
|
||||
帧 = ("data: " + json.dumps(完成事件(), ensure_ascii=False) + "\r\n\r\n").encode()
|
||||
|
||||
async def 字节流():
|
||||
for b in 帧:
|
||||
yield bytes([b])
|
||||
|
||||
async def 收集():
|
||||
return [项 async for 项 in 解码SSE(字节流())]
|
||||
|
||||
结果 = 解析响应(asyncio.run(收集()))
|
||||
assert 结果.状态 == "completed" and 结果.文本 == "中文\u0085保持。"
|
||||
assert 结果.实际模型 == "test-model"
|
||||
assert 结果.用量.输入token == 10 and 结果.用量.输出token == 4
|
||||
|
||||
|
||||
@pytest.mark.parametrize(
|
||||
"事件,期望",
|
||||
[
|
||||
([{"type": "response.output_text.delta", "delta": "只到一半"}], "incomplete"),
|
||||
([完成事件(status="incomplete")], "incomplete"),
|
||||
([完成事件(status="failed")], "failed"),
|
||||
([{"type": "error", "code": "server_error"}], "failed"),
|
||||
],
|
||||
ids=["no-terminal", "incomplete", "failed", "error"],
|
||||
)
|
||||
def test_断流与失败不计模型完成__a61002(事件, 期望) -> None:
|
||||
assert 解析响应(事件).状态 == 期望
|
||||
|
||||
|
||||
def test_完成但用量未知独立保留__a61003() -> None:
|
||||
结果 = 解析响应([完成事件(usage=False)])
|
||||
assert 结果.状态 == "completed" and 结果.用量 is None
|
||||
|
||||
|
||||
def test_错误SSE帧和UTF8拒绝__a61004() -> None:
|
||||
async def 字节流():
|
||||
yield b'data: {"text":"\xff"}\n\n'
|
||||
|
||||
async def 收集():
|
||||
return [项 async for 项 in 解码SSE(字节流())]
|
||||
|
||||
with pytest.raises(模型协议错误):
|
||||
asyncio.run(收集())
|
||||
|
||||
|
||||
@pytest.mark.parametrize(
|
||||
"结束,原因,状态",
|
||||
[
|
||||
(True, "end_turn", "completed"),
|
||||
(False, "end_turn", "incomplete"),
|
||||
(True, "max_tokens", "incomplete"),
|
||||
],
|
||||
ids=["complete", "missing-stop", "token-limit"],
|
||||
)
|
||||
def test_Anthropic结束消息与停止原因共同判定__a61005(结束, 原因, 状态) -> None:
|
||||
from muse.基础设施.模型.Messages import 解析响应 as 解析
|
||||
|
||||
事件 = [
|
||||
{
|
||||
"type": "message_start",
|
||||
"message": {"id": "a1", "model": "opus", "usage": {"input_tokens": 3}},
|
||||
},
|
||||
{"type": "content_block_start", "index": 0, "content_block": {"type": "text", "text": ""}},
|
||||
{
|
||||
"type": "content_block_delta",
|
||||
"index": 0,
|
||||
"delta": {"type": "text_delta", "text": "甲乙。"},
|
||||
},
|
||||
{"type": "message_delta", "delta": {"stop_reason": 原因}, "usage": {"output_tokens": 2}},
|
||||
]
|
||||
if 结束:
|
||||
事件.append({"type": "message_stop"})
|
||||
结果 = 解析(事件)
|
||||
assert 结果.状态 == 状态 and 结果.文本 == "甲乙。"
|
||||
assert 结果.用量.输入token == 3
|
||||
|
||||
|
||||
def test_ChatCompletions完成与工具回合分开__a61006() -> None:
|
||||
from muse.基础设施.模型.ChatCompletions import 解析响应 as 解析
|
||||
|
||||
结果 = 解析(
|
||||
[
|
||||
{
|
||||
"id": "c1",
|
||||
"model": "test",
|
||||
"choices": [{"index": 0, "delta": {"content": "完成。"}, "finish_reason": "stop"}],
|
||||
}
|
||||
]
|
||||
)
|
||||
assert 结果.状态 == "completed" and 结果.用量 is None
|
||||
工具 = 解析(
|
||||
[
|
||||
{
|
||||
"id": "c2",
|
||||
"model": "test",
|
||||
"choices": [
|
||||
{
|
||||
"index": 0,
|
||||
"delta": {
|
||||
"tool_calls": [
|
||||
{
|
||||
"index": 0,
|
||||
"id": "tool1",
|
||||
"function": {"name": "read", "arguments": "{}"},
|
||||
}
|
||||
]
|
||||
},
|
||||
"finish_reason": "tool_calls",
|
||||
}
|
||||
],
|
||||
}
|
||||
]
|
||||
)
|
||||
assert 工具.状态 == "tool_calls" and 工具.工具调用[0].名称 == "read"
|
||||
|
||||
|
||||
@pytest.mark.parametrize(
|
||||
"模型,文本,合法",
|
||||
[
|
||||
("test", '{"正文":"中文"}', True),
|
||||
("other", '{"正文":"中文"}', False),
|
||||
("test", '{"正文":1}', False),
|
||||
("test", '{"正文":"中文","state":"confirmed"}', False),
|
||||
],
|
||||
ids=["valid", "model-drift", "wrong-type", "extra-authority"],
|
||||
)
|
||||
def test_实际模型与输出合同同时满足才产候选__a61007(模型, 文本, 合法) -> None:
|
||||
from muse.任务运行.接口 import 校验模型输出, 模型结果, 模型请求
|
||||
|
||||
请求 = 模型请求(
|
||||
"call",
|
||||
"synthetic",
|
||||
"test",
|
||||
"合成系统说明",
|
||||
"合成输入",
|
||||
{
|
||||
"type": "object",
|
||||
"properties": {"正文": {"type": "string"}},
|
||||
"required": ["正文"],
|
||||
"additionalProperties": False,
|
||||
},
|
||||
100,
|
||||
30,
|
||||
)
|
||||
结果 = 模型结果("completed", 文本, 模型, None)
|
||||
if 合法:
|
||||
assert 校验模型输出(请求, 结果) == {"正文": "中文"}
|
||||
else:
|
||||
with pytest.raises(模型协议错误):
|
||||
校验模型输出(请求, 结果)
|
||||
66
tests/契约/test_角色能力与配置.py
Normal file
66
tests/契约/test_角色能力与配置.py
Normal file
@ -0,0 +1,66 @@
|
||||
"""角色策略来自可审配置;不发模型请求。"""
|
||||
|
||||
from pathlib import Path
|
||||
|
||||
import pytest
|
||||
import yaml
|
||||
|
||||
from muse.任务运行.接口 import 模型协议错误, 角色策略目录
|
||||
|
||||
仓库根 = Path(__file__).resolve().parents[2]
|
||||
源码根 = 仓库根 / "src" / "muse"
|
||||
|
||||
|
||||
@pytest.mark.parametrize("角色", ["writer", "planner", "judge"], ids=["writer", "planner", "judge"])
|
||||
def test_固定角色拒绝降低模型能力__a63001(角色) -> None:
|
||||
目录 = 角色策略目录(
|
||||
yaml.safe_load((Path(__file__).resolve().parents[2] / "配置/角色策略.yaml").read_text())
|
||||
)
|
||||
with pytest.raises(模型协议错误):
|
||||
目录.冻结(角色, provider="configured", model="deepseek-v4-flash", 阶段="生成")
|
||||
固定 = 目录.冻结(
|
||||
角色, provider="configured", model="claude-opus-4-8[1M]", thinking="high", 阶段="生成"
|
||||
)
|
||||
assert 固定.允许实际模型 == ("claude-opus-4-8[1M]", "claude-opus-4-8")
|
||||
|
||||
|
||||
def test_写手生成与盲评拒绝工具__a63002() -> None:
|
||||
目录 = 角色策略目录(
|
||||
yaml.safe_load((Path(__file__).resolve().parents[2] / "配置/角色策略.yaml").read_text())
|
||||
)
|
||||
for 角色, 阶段 in [("writer", "生成"), ("blind_judge", "执行")]:
|
||||
with pytest.raises(模型协议错误):
|
||||
目录.冻结(
|
||||
角色, provider="configured", model="claude-opus-4-8[1M]", 工具=("read",), 阶段=阶段
|
||||
)
|
||||
assert (
|
||||
目录.冻结("semantic_detector", provider="configured", model="MiniMax-M3").角色 == "detector"
|
||||
)
|
||||
|
||||
|
||||
def test_active_runtime_has_no_embedded_api_token__5274c6() -> None:
|
||||
"""活跃运行代码与配置模板不得内嵌凭据;密钥只能来自受控存储引用。"""
|
||||
import re
|
||||
|
||||
模式 = [
|
||||
re.compile(r"sk-[A-Za-z0-9]{20,}"),
|
||||
re.compile(
|
||||
r"(?:api[-_]?key|token|secret|password)\s*[:=]\s*[\"'][^\"{}$\n]{8,}[\"']", re.I
|
||||
),
|
||||
]
|
||||
违规 = []
|
||||
for 文件 in 源码根.rglob("*.py"):
|
||||
if "__pycache__" in 文件.parts:
|
||||
continue
|
||||
文本 = 文件.read_text(encoding="utf-8")
|
||||
for 规则 in 模式:
|
||||
命中 = 规则.search(文本)
|
||||
if 命中:
|
||||
违规.append(f"{文件.relative_to(仓库根)}: {命中.group(0)[:24]}…")
|
||||
for 文件 in (仓库根 / "配置").glob("*.toml"):
|
||||
文本 = 文件.read_text(encoding="utf-8")
|
||||
for 规则 in 模式:
|
||||
命中 = 规则.search(文本)
|
||||
if 命中:
|
||||
违规.append(f"配置/{文件.name}: {命中.group(0)[:24]}…")
|
||||
assert 违规 == [], "发现内嵌凭据形状:\n" + "\n".join(违规)
|
||||
888
tests/集成/test_受控模型调用.py
Normal file
888
tests/集成/test_受控模型调用.py
Normal file
@ -0,0 +1,888 @@
|
||||
"""真实任务和PG账本贯通合成HTTP流;不据此宣称真实模型或文学效果通过。"""
|
||||
|
||||
import asyncio
|
||||
import hashlib
|
||||
import json
|
||||
from datetime import UTC, datetime, timedelta
|
||||
from decimal import Decimal
|
||||
from http.server import BaseHTTPRequestHandler, ThreadingHTTPServer
|
||||
from threading import Thread
|
||||
|
||||
import pytest
|
||||
|
||||
from muse.任务运行.接口 import (
|
||||
任务服务,
|
||||
任务状态,
|
||||
任务请求,
|
||||
原文服务,
|
||||
原文错误,
|
||||
执行上下文,
|
||||
模型协议错误,
|
||||
模型执行器,
|
||||
模型请求,
|
||||
步骤处理器,
|
||||
步骤结果,
|
||||
步骤计划,
|
||||
角色策略目录,
|
||||
证据服务,
|
||||
请求字节,
|
||||
)
|
||||
from muse.任务运行.预算管理 import (
|
||||
任务预算计划,
|
||||
角色预算,
|
||||
预算不足,
|
||||
预算状态冲突,
|
||||
预算管理,
|
||||
额度策略,
|
||||
)
|
||||
from muse.共享.调用身份 import 内容用途, 用途
|
||||
from muse.基础设施.受控文件 import 受控文件
|
||||
from muse.基础设施.宿主.直接调用 import 直接宿主
|
||||
from muse.基础设施.模型.HTTP传输 import HTTP传输
|
||||
from muse.编排.接口 import 流程定义, 流程服务, 流程登记
|
||||
|
||||
pytestmark = pytest.mark.数据库
|
||||
|
||||
|
||||
@pytest.fixture
|
||||
def 合成HTTP():
|
||||
服务, 异常 = [], []
|
||||
|
||||
def 启动(处理):
|
||||
class 请求处理(BaseHTTPRequestHandler):
|
||||
def do_POST(self):
|
||||
try:
|
||||
输入 = json.loads(self.rfile.read(int(self.headers["Content-Length"])))
|
||||
输出 = 处理(输入)
|
||||
self.send_response(200)
|
||||
self.send_header("Content-Type", "text/event-stream")
|
||||
self.send_header("Content-Length", str(len(输出)))
|
||||
self.end_headers()
|
||||
self.wfile.write(输出)
|
||||
except Exception as exc:
|
||||
异常.append(exc)
|
||||
self.send_error(500)
|
||||
|
||||
def log_message(self, *args):
|
||||
pass
|
||||
|
||||
server = ThreadingHTTPServer(("127.0.0.1", 0), 请求处理)
|
||||
thread = Thread(target=server.serve_forever, daemon=True)
|
||||
thread.start()
|
||||
服务.append((server, thread))
|
||||
return f"http://127.0.0.1:{server.server_port}/v1/responses"
|
||||
|
||||
yield 启动
|
||||
for server, thread in 服务:
|
||||
server.shutdown()
|
||||
server.server_close()
|
||||
thread.join(timeout=2)
|
||||
assert not 异常
|
||||
|
||||
|
||||
class 合成计价:
|
||||
版本 = "synthetic-price-1"
|
||||
|
||||
def 金额(self, 结果):
|
||||
return Decimal("0.125") if 结果.用量 else None
|
||||
|
||||
|
||||
@pytest.mark.parametrize("场景", ["temporary", "insufficient", "write-failure"])
|
||||
def test_临时模型原文遵守租约且普通证据不保留字节__a82110(模型环境, tmp_path, 合成HTTP, 场景):
|
||||
from dataclasses import replace
|
||||
|
||||
库, 运行, 策略, 初始, 原文, _, 预算, 上下文 = 模型环境
|
||||
请求 = replace(初始, 用户输入="仅在租约内可用的合成输入")
|
||||
输入哈希 = hashlib.sha256(请求字节(请求)).hexdigest()
|
||||
授权 = 原文.批准保留(
|
||||
上下文.领取.任务ID,
|
||||
"author",
|
||||
"temporary-approve",
|
||||
来源版本="temporary-source-1",
|
||||
哈希=(输入哈希,),
|
||||
内容用途="generation",
|
||||
方式="temporary",
|
||||
有效期=datetime.now(UTC) + timedelta(minutes=5),
|
||||
调用ID=请求.调用ID,
|
||||
调用请求哈希=输入哈希,
|
||||
)
|
||||
租约 = 原文.创建租约(
|
||||
上下文.领取.任务ID,
|
||||
授权,
|
||||
"temporary-lease",
|
||||
datetime.now(UTC) + timedelta(seconds=2 if 场景 == "insufficient" else 120),
|
||||
最少剩余秒=0,
|
||||
)
|
||||
收到 = []
|
||||
|
||||
def 回应(数据):
|
||||
收到.append(数据)
|
||||
assert 数据["input"] == 请求.用户输入
|
||||
if 场景 == "write-failure":
|
||||
# 输入已写入;模拟输出返回前此租约目录实际失去合法写入权限。
|
||||
(原文.文件.根 / 租约).chmod(0o500)
|
||||
return (
|
||||
"data: "
|
||||
+ json.dumps(
|
||||
{
|
||||
"type": "response.completed",
|
||||
"response": {
|
||||
"id": "temporary-response",
|
||||
"model": "claude-opus-4-8",
|
||||
"status": "completed",
|
||||
"output": [
|
||||
{
|
||||
"type": "message",
|
||||
"content": [
|
||||
{"type": "output_text", "text": '{"text":"租约内响应"}'}
|
||||
],
|
||||
}
|
||||
],
|
||||
"usage": {"input_tokens": 1, "output_tokens": 1},
|
||||
},
|
||||
}
|
||||
)
|
||||
+ "\n\n"
|
||||
).encode()
|
||||
|
||||
凭据 = tmp_path / "temporary-provider-key"
|
||||
凭据.write_text("synthetic-temporary-provider")
|
||||
执行器 = 模型执行器(
|
||||
运行,
|
||||
预算,
|
||||
原文,
|
||||
证据服务(库),
|
||||
策略,
|
||||
直接宿主(HTTP传输(合成HTTP(回应), "受控存储", str(凭据), "responses")),
|
||||
合成计价(),
|
||||
)
|
||||
if 场景 == "insufficient":
|
||||
with pytest.raises(原文错误, match="余量|期限"):
|
||||
asyncio.run(执行器.执行(上下文, 请求, 阶段="生成", 原文授权ID=授权, 临时租约ID=租约))
|
||||
assert not 收到
|
||||
with 库.连接() as 连:
|
||||
assert 连.execute("SELECT count(*) FROM muse_budget_reservation").fetchone()[0] == 0
|
||||
assert 连.execute("SELECT count(*) FROM muse_runtime_evidence").fetchone()[0] == 0
|
||||
elif 场景 == "write-failure":
|
||||
from muse.基础设施.受控文件 import 受控文件错误
|
||||
|
||||
try:
|
||||
with pytest.raises(受控文件错误):
|
||||
asyncio.run(
|
||||
执行器.执行(上下文, 请求, 阶段="生成", 原文授权ID=授权, 临时租约ID=租约)
|
||||
)
|
||||
finally:
|
||||
(原文.文件.根 / 租约).chmod(0o700)
|
||||
assert len(收到) == 1 and 预算.读取(请求.调用ID).实际金额 == Decimal("0.125")
|
||||
with 库.连接() as 连:
|
||||
assert 连.execute(
|
||||
"SELECT outcome,content,metadata->>'error_code' FROM muse_runtime_evidence "
|
||||
"WHERE kind='failure'"
|
||||
).fetchone() == ("failed", None, "RAW_TEMP_WRITE_FAILED")
|
||||
运行.失败步骤(上下文.领取, "raw_retention_failed")
|
||||
assert 运行.读取任务(上下文.领取.任务ID).状态 == 任务状态.待对账
|
||||
assert 原文.读取(上下文.领取.任务ID, 租约, 输入哈希) == 请求字节(请求)
|
||||
else:
|
||||
交付 = asyncio.run(执行器.执行(上下文, 请求, 阶段="生成", 原文授权ID=授权, 临时租约ID=租约))
|
||||
assert 交付.内容 == {"text": "租约内响应"}
|
||||
assert 交付.证据回执["retention"] == "hash_only"
|
||||
assert 交付.证据回执["raw_lease_id"] == 租约
|
||||
assert 原文.读取(上下文.领取.任务ID, 租约, 输入哈希) == 请求字节(请求)
|
||||
响应 = json.loads(原文.读取(上下文.领取.任务ID, 租约, 交付.证据回执["content_hash"]))
|
||||
assert 响应["文本"] == '{"text":"租约内响应"}'
|
||||
with 库.连接() as 连:
|
||||
assert (
|
||||
连.execute(
|
||||
"SELECT count(*) FROM muse_runtime_evidence WHERE content IS NOT NULL"
|
||||
).fetchone()[0]
|
||||
== 0
|
||||
)
|
||||
assert 预算.读取(请求.调用ID).实际金额 == Decimal("0.125")
|
||||
原文.清理(上下文.领取.任务ID, 租约, 作者="author")
|
||||
assert 原文.状态(上下文.领取.任务ID, 租约)["state"] == "closed"
|
||||
with pytest.raises(原文错误):
|
||||
原文.读取(上下文.领取.任务ID, 租约, 交付.证据回执["content_hash"])
|
||||
assert (
|
||||
证据服务(库).读取回执(上下文.领取.任务ID, 交付.证据回执["evidence_id"]) == 交付.证据回执
|
||||
)
|
||||
|
||||
|
||||
@pytest.mark.parametrize("模型环境", ["configured"], indirect=True, ids=["evaluation"])
|
||||
def test_任务绑定配置装配真实传输且不随新版本切换__a62007(模型环境, tmp_path, 合成HTTP):
|
||||
"""配置验证器为明确离线替身;验证评测配置到真实HTTP装配,不声明生产验证通过。"""
|
||||
from dataclasses import replace
|
||||
|
||||
from muse.任务运行.模型 import 内容哈希
|
||||
from muse.任务运行.配置版本 import (
|
||||
凭据引用,
|
||||
提供方配置,
|
||||
运行配置内容,
|
||||
配置版本管理,
|
||||
配置验证证据,
|
||||
)
|
||||
from muse.启动 import 构建
|
||||
from muse.配置 import 应用配置
|
||||
|
||||
库, 运行, 策略, 请求, 原文, 授权, 预算, 上下文 = 模型环境
|
||||
收到 = []
|
||||
|
||||
def 回应(值):
|
||||
收到.append(值)
|
||||
return (
|
||||
"data: "
|
||||
+ json.dumps(
|
||||
{
|
||||
"type": "response.completed",
|
||||
"response": {
|
||||
"id": "configured-response",
|
||||
"model": "claude-opus-4-8",
|
||||
"status": "completed",
|
||||
"output": [
|
||||
{
|
||||
"type": "message",
|
||||
"content": [{"type": "output_text", "text": '{"text":"固定配置"}'}],
|
||||
}
|
||||
],
|
||||
"usage": {"input_tokens": 1, "output_tokens": 1},
|
||||
},
|
||||
}
|
||||
)
|
||||
+ "\n\n"
|
||||
).encode()
|
||||
|
||||
凭据 = tmp_path / "configured-key"
|
||||
凭据.write_text("synthetic-configured-credential")
|
||||
原配置 = 运行配置内容(
|
||||
"direct",
|
||||
"1",
|
||||
策略.定义["version"],
|
||||
策略.资源发布身份,
|
||||
"synthetic",
|
||||
{"writer": {"provider": "synthetic", "model": 请求.model, "thinking": "high", "tools": []}},
|
||||
(凭据引用("provider-key", "受控存储", str(凭据)),),
|
||||
(提供方配置("synthetic", "responses", 合成HTTP(回应), "provider-key"),),
|
||||
"synthetic-price-1",
|
||||
)
|
||||
|
||||
class 离线配置检查:
|
||||
身份 = "synthetic-config-validation"
|
||||
|
||||
def 验证(self, 内容, 执行用途):
|
||||
return 配置验证证据(
|
||||
内容哈希(内容.冻结()),
|
||||
内容.角色策略版本,
|
||||
内容.资源发布身份,
|
||||
执行用途,
|
||||
("synthetic-configuration-contract",),
|
||||
"offline_contract",
|
||||
)
|
||||
|
||||
管理 = 配置版本管理(库, 离线配置检查())
|
||||
管理.保存草案("model-config", "1", 原配置)
|
||||
回执 = 管理.验证版本("model-config", "1")
|
||||
管理.启用("model-config", "1", 验证回执=回执, 批准引用="synthetic-approval", 预期代次=0)
|
||||
管理.冻结到任务(上下文.领取.任务ID, "model-config")
|
||||
# 新配置指向不可用地址;旧任务仍须消费其原版本。
|
||||
管理.保存草案(
|
||||
"model-config",
|
||||
"2",
|
||||
replace(
|
||||
原配置,
|
||||
提供方=(
|
||||
提供方配置("synthetic", "responses", "http://127.0.0.1:1/unused", "provider-key"),
|
||||
),
|
||||
),
|
||||
)
|
||||
R2 = 管理.验证版本("model-config", "2")
|
||||
管理.启用("model-config", "2", 验证回执=R2, 批准引用="synthetic-approval-2", 预期代次=1)
|
||||
装配 = 构建(
|
||||
应用配置(库.引用, 策略.资源发布身份, 运行用途=用途.评测),
|
||||
流程=运行.处理器,
|
||||
)
|
||||
执行器 = 装配.要求模型执行器(上下文, 合成计价())
|
||||
结果 = asyncio.run(执行器.执行(上下文, 请求, 阶段="生成", 原文授权ID=授权))
|
||||
assert 结果.内容 == {"text": "固定配置"}
|
||||
assert len(收到) == 1 and 收到[0]["reasoning"] == {"effort": "high"}
|
||||
assert 预算.读取(请求.调用ID).实际金额 == Decimal("0.125")
|
||||
assert 管理.读取任务绑定(上下文.领取.任务ID).版本 == "1"
|
||||
with 库.连接() as 连:
|
||||
assert (
|
||||
连.execute(
|
||||
"SELECT metadata->>'runtime_config_hash' FROM evaluation.muse_runtime_evidence "
|
||||
"WHERE kind='model_response'"
|
||||
).fetchone()[0]
|
||||
== 管理.读取任务绑定(上下文.领取.任务ID).内容哈希
|
||||
)
|
||||
with pytest.raises(模型协议错误, match="配置"):
|
||||
asyncio.run(
|
||||
执行器.执行(上下文, replace(请求, provider="other"), 阶段="生成", 原文授权ID=授权)
|
||||
)
|
||||
assert len(收到) == 1
|
||||
|
||||
|
||||
def test_CLI保存查看配置草案保留引用且不自动启用__a62008(应用测试库, tmp_path):
|
||||
import subprocess
|
||||
import sys
|
||||
from pathlib import Path
|
||||
|
||||
from muse.任务运行.配置版本 import 配置版本管理
|
||||
from muse.共享.错误 import Muse错误
|
||||
|
||||
库 = 应用测试库[用途.评测]
|
||||
应用文件 = tmp_path / "app.toml"
|
||||
应用文件.write_text(
|
||||
'["数据库"]\n"取值方式"="受控存储"\n"位置"='
|
||||
+ json.dumps(库.引用.位置, ensure_ascii=False)
|
||||
+ '\n["资源"]\n"发布身份"="synthetic"\n["运行"]\n"用途"="evaluation"\n'
|
||||
)
|
||||
内容 = Path(__file__).parents[2] / "配置/提供方.example.toml"
|
||||
基础命令 = [sys.executable, "-I", "-m", "muse", "配置", str(应用文件)]
|
||||
保存 = subprocess.run(
|
||||
基础命令 + ["保存", "cli-model", "1", "--内容", str(内容)],
|
||||
cwd=tmp_path,
|
||||
capture_output=True,
|
||||
text=True,
|
||||
)
|
||||
assert 保存.returncode == 0, 保存.stderr
|
||||
查看 = subprocess.run(
|
||||
基础命令 + ["查看", "cli-model", "1"], cwd=tmp_path, capture_output=True, text=True
|
||||
)
|
||||
assert 查看.returncode == 0, 查看.stderr
|
||||
assert json.loads(保存.stdout) == json.loads(查看.stdout)
|
||||
assert json.loads(查看.stdout)["内容"]["凭据"][0]["来源"] == "受控存储"
|
||||
with 库.连接() as 连:
|
||||
assert (
|
||||
连.execute("SELECT count(*) FROM evaluation.muse_runtime_config_active").fetchone()[0]
|
||||
== 0
|
||||
)
|
||||
# 不注入验证器时,读取和保存可用,验证不得自行产生成功回执。
|
||||
with pytest.raises(Muse错误, match="验证器"):
|
||||
配置版本管理(库).验证版本("cli-model", "1")
|
||||
|
||||
|
||||
@pytest.fixture
|
||||
def 模型环境(应用测试库, tmp_path, request):
|
||||
次数 = getattr(request, "param", 1)
|
||||
评测配置 = 次数 == "configured"
|
||||
有工具 = 次数 == "session"
|
||||
次数 = 2 if 有工具 else 1 if 评测配置 else 次数
|
||||
角色 = "planner" if 有工具 else "writer"
|
||||
工具 = ("schema_read",) if 有工具 else ()
|
||||
库 = 应用测试库[用途.评测 if 评测配置 else 用途.生产]
|
||||
策略 = 角色策略目录.从发布包()
|
||||
登记 = 流程登记()
|
||||
登记.登记处理器(
|
||||
步骤处理器("model-probe", "1", lambda _: 步骤结果({}), "1", "1", 角色=角色, 允许工具=工具)
|
||||
)
|
||||
登记.登记类型("model-probe", 必需保护=())
|
||||
|
||||
def 重验合成输入(快照):
|
||||
# 合成任务不读取业务正文,来源范围为空;固定结构版本由只读工具显式读取。
|
||||
assert 快照.冻结输入["冻结上下文"]["source_scope"] == {}
|
||||
assert 快照.冻结输入["资源发布身份"] == 策略.资源发布身份
|
||||
|
||||
登记.登记恢复检查("model-probe", "1", 重验合成输入)
|
||||
运行 = 任务服务(库, 登记)
|
||||
流程 = 流程服务(运行, 登记)
|
||||
流程.发布(
|
||||
流程定义("model-probe", "1", (步骤计划("call", "model-probe", "1", 角色=角色, 工具=工具),))
|
||||
)
|
||||
任务 = 流程.发起(
|
||||
任务请求(
|
||||
"model-probe",
|
||||
"command",
|
||||
"author",
|
||||
库.用途,
|
||||
内容用途.生成,
|
||||
{},
|
||||
策略.定义["version"],
|
||||
策略.资源发布身份,
|
||||
{
|
||||
"source_scope": {},
|
||||
"schema_versions": {},
|
||||
"authorization": "grant",
|
||||
"budget": {},
|
||||
"stop_conditions": [],
|
||||
},
|
||||
),
|
||||
"model-probe",
|
||||
"1",
|
||||
)
|
||||
领取 = 运行.领取步骤("worker", ["model-probe"])
|
||||
assert 领取 is not None
|
||||
请求 = 模型请求(
|
||||
"call-1",
|
||||
"synthetic",
|
||||
"claude-opus-4-8[1M]",
|
||||
"返回合成测试JSON",
|
||||
"一次合成调用",
|
||||
{
|
||||
"type": "object",
|
||||
"required": ["text"],
|
||||
"properties": {"text": {"type": "string"}},
|
||||
"additionalProperties": False,
|
||||
},
|
||||
64,
|
||||
10,
|
||||
thinking="high" if 评测配置 else None,
|
||||
允许实际模型=("claude-opus-4-8",),
|
||||
)
|
||||
原文 = 原文服务(库, 受控文件(tmp_path / "raw"))
|
||||
哈希 = hashlib.sha256(请求字节(请求)).hexdigest()
|
||||
授权 = 原文.批准保留(
|
||||
任务,
|
||||
"author",
|
||||
"approve",
|
||||
来源版本="synthetic-1",
|
||||
哈希=(哈希,),
|
||||
内容用途="generation",
|
||||
方式="persistent",
|
||||
有效期=datetime.now(UTC) + timedelta(minutes=5),
|
||||
调用ID=请求.调用ID,
|
||||
调用请求哈希=哈希,
|
||||
)
|
||||
预算 = 预算管理(库, "synthetic")
|
||||
预算.登记策略(额度策略("synthetic", "1", Decimal("2"), 10))
|
||||
预算.登记任务预算(
|
||||
任务,
|
||||
任务预算计划(
|
||||
Decimal(str(次数)),
|
||||
(
|
||||
角色预算(
|
||||
角色,
|
||||
次数,
|
||||
次数,
|
||||
Decimal("1"),
|
||||
),
|
||||
),
|
||||
"approval",
|
||||
datetime.now(UTC) + timedelta(minutes=5),
|
||||
),
|
||||
)
|
||||
return 库, 运行, 策略, 请求, 原文, 授权, 预算, 执行上下文(领取, 运行.读取任务(任务))
|
||||
|
||||
|
||||
@pytest.mark.parametrize("模型环境", [2], indirect=True, ids=["two-calls"])
|
||||
def test_同尝试两个相同回复各自保留调用证据__a62005(模型环境, tmp_path, 合成HTTP):
|
||||
from dataclasses import replace
|
||||
|
||||
库, 运行, 策略, 请求, 原文, 授权, 预算, 上下文 = 模型环境
|
||||
已调用 = []
|
||||
|
||||
def 回应(value):
|
||||
已调用.append(value)
|
||||
数据 = {
|
||||
"type": "response.completed",
|
||||
"response": {
|
||||
"id": f"reply-{len(已调用)}",
|
||||
"model": "claude-opus-4-8",
|
||||
"status": "completed",
|
||||
"output": [
|
||||
{
|
||||
"type": "message",
|
||||
"content": [{"type": "output_text", "text": '{"text":"同一个结果"}'}],
|
||||
}
|
||||
],
|
||||
"usage": {"input_tokens": 1, "output_tokens": 1},
|
||||
},
|
||||
}
|
||||
return ("data: " + json.dumps(数据) + "\n\n").encode()
|
||||
|
||||
凭据 = tmp_path / "same-output-key"
|
||||
凭据.write_text("synthetic-credential")
|
||||
执行器 = 模型执行器(
|
||||
运行,
|
||||
预算,
|
||||
原文,
|
||||
证据服务(库),
|
||||
策略,
|
||||
直接宿主(HTTP传输(合成HTTP(回应), "受控存储", str(凭据), "responses")),
|
||||
合成计价(),
|
||||
)
|
||||
首次 = asyncio.run(执行器.执行(上下文, 请求, 阶段="生成", 原文授权ID=授权))
|
||||
第二请求 = replace(请求, 调用ID="call-2")
|
||||
请求哈希 = hashlib.sha256(请求字节(第二请求)).hexdigest()
|
||||
第二授权 = 原文.批准保留(
|
||||
上下文.领取.任务ID,
|
||||
"author",
|
||||
"approve-second",
|
||||
来源版本="synthetic-1",
|
||||
哈希=(请求哈希,),
|
||||
内容用途="generation",
|
||||
方式="persistent",
|
||||
有效期=datetime.now(UTC) + timedelta(minutes=5),
|
||||
调用ID=第二请求.调用ID,
|
||||
调用请求哈希=请求哈希,
|
||||
)
|
||||
第二次 = asyncio.run(执行器.执行(上下文, 第二请求, 阶段="生成", 原文授权ID=第二授权))
|
||||
assert 首次.内容 == 第二次.内容 and len(已调用) == 2
|
||||
assert 首次.证据回执["evidence_id"] != 第二次.证据回执["evidence_id"]
|
||||
assert 预算.读取("call-1").实际金额 == 预算.读取("call-2").实际金额 == Decimal("0.125")
|
||||
|
||||
|
||||
@pytest.mark.parametrize("模型环境", ["session"], indirect=True, ids=["session"])
|
||||
@pytest.mark.parametrize(
|
||||
"宿主类型",
|
||||
[
|
||||
"direct",
|
||||
pytest.param("pi", marks=pytest.mark.宿主),
|
||||
pytest.param("resume", marks=pytest.mark.宿主),
|
||||
pytest.param("reconcile", marks=pytest.mark.宿主),
|
||||
],
|
||||
)
|
||||
def test_有限角色会话贯通真实工具与逐回合账本__a62006(
|
||||
模型环境, 应用测试库, tmp_path, 合成HTTP, 宿主类型
|
||||
):
|
||||
import os
|
||||
from dataclasses import replace
|
||||
from pathlib import Path
|
||||
|
||||
from muse.任务运行.接口 import 工具定义, 工具结果
|
||||
from muse.任务运行.角色会话 import 组装角色请求, 角色会话
|
||||
from muse.元数据.接口 import 元数据服务, 导入内置结构
|
||||
from muse.基础设施.宿主.直接调用 import 直接角色循环
|
||||
|
||||
库, 运行, 策略, 初始, 原文, _, 预算, 上下文 = 模型环境
|
||||
with 应用测试库[用途.维护].连接() as 连, 连.transaction():
|
||||
导入内置结构(元数据服务(连))
|
||||
读取次数 = []
|
||||
|
||||
def 读取结构(范围, 参数):
|
||||
读取次数.append(范围.任务ID)
|
||||
with 库.连接() as 连:
|
||||
结构 = 元数据服务(连).读取结构("work_core", 1)
|
||||
return 工具结果({"schema_id": 结构.schema_id, "schema_hash": 结构.内容哈希}, ())
|
||||
|
||||
定义 = 工具定义(
|
||||
"schema_read",
|
||||
"读取已发布作品结构",
|
||||
{"type": "object", "properties": {}, "additionalProperties": False},
|
||||
读取结构,
|
||||
)
|
||||
请求 = 组装角色请求(
|
||||
策略,
|
||||
"planner",
|
||||
replace(
|
||||
初始,
|
||||
工具=(
|
||||
{
|
||||
"name": 定义.名称,
|
||||
"description": 定义.说明,
|
||||
"parameters": 定义.参数合同,
|
||||
},
|
||||
),
|
||||
),
|
||||
)
|
||||
from fastapi.testclient import TestClient
|
||||
|
||||
from muse.接入.http.应用 import 创建应用
|
||||
from muse.配置 import 应用配置, 服务配置
|
||||
|
||||
口令 = tmp_path / "session-author-key"
|
||||
口令.write_text("synthetic-session-access")
|
||||
配置 = 应用配置(
|
||||
库.引用,
|
||||
"test",
|
||||
HTTP=服务配置(
|
||||
str(口令),
|
||||
作者ID="author",
|
||||
公开地址="http://testserver",
|
||||
允许来源=("http://testserver",),
|
||||
),
|
||||
原文暂存=str(tmp_path / "session-http-raw"),
|
||||
)
|
||||
路径 = f"/api/v1/tasks/{上下文.领取.任务ID}/role-session-authorizations"
|
||||
批准请求 = {
|
||||
"command_id": "approve-session",
|
||||
"step_id": "call",
|
||||
"stage": "执行",
|
||||
"initial_request_hash": hashlib.sha256(请求字节(请求)).hexdigest(),
|
||||
"max_model_calls": 2,
|
||||
"max_tool_calls": 1,
|
||||
"valid_until": (datetime.now(UTC) + timedelta(minutes=5)).isoformat(),
|
||||
}
|
||||
with TestClient(创建应用(配置)) as client:
|
||||
client.headers["Origin"] = "http://testserver"
|
||||
assert client.post(路径, json=批准请求).status_code == 401
|
||||
assert (
|
||||
client.post(
|
||||
"/api/v1/session", json={"password": "synthetic-session-access"}
|
||||
).status_code
|
||||
== 200
|
||||
)
|
||||
assert client.post(路径, json={**批准请求, "approved_by": "other"}).status_code == 422
|
||||
批准结果 = client.post(路径, json=批准请求)
|
||||
assert 批准结果.status_code == 200, 批准结果.text
|
||||
会话授权 = 批准结果.json()["authorization_id"]
|
||||
收到 = []
|
||||
|
||||
def 回应(数据):
|
||||
收到.append(数据)
|
||||
assert 数据["instructions"] == 请求.系统提示
|
||||
if len(收到) == 1:
|
||||
输出 = [
|
||||
{
|
||||
"type": "function_call",
|
||||
"id": "fc-1",
|
||||
"call_id": "schema-call",
|
||||
"name": "schema_read",
|
||||
"arguments": "{}",
|
||||
}
|
||||
]
|
||||
else:
|
||||
assert 数据["input"][1]["call_id"] == "schema-call"
|
||||
工具数据 = json.loads(数据["input"][2]["output"])
|
||||
assert 工具数据["内容"]["schema_id"] == "work_core"
|
||||
输出 = [
|
||||
{
|
||||
"type": "message",
|
||||
"content": [
|
||||
{"type": "output_text", "text": '{"text":"结构已读取,规划待作者确认"}'}
|
||||
],
|
||||
}
|
||||
]
|
||||
数据 = {
|
||||
"type": "response.completed",
|
||||
"response": {
|
||||
"id": f"session-response-{len(收到)}",
|
||||
"model": "claude-opus-4-8",
|
||||
"status": "completed",
|
||||
"output": 输出,
|
||||
"usage": {"input_tokens": 3, "output_tokens": 4},
|
||||
},
|
||||
}
|
||||
return ("data: " + json.dumps(数据, ensure_ascii=False) + "\n\n").encode()
|
||||
|
||||
凭据 = tmp_path / "session-key"
|
||||
凭据.write_text("synthetic-credential")
|
||||
|
||||
class 会话计价(合成计价):
|
||||
未知首次 = 宿主类型 == "reconcile"
|
||||
|
||||
def 金额(self, 结果):
|
||||
if self.未知首次:
|
||||
self.未知首次 = False
|
||||
return None
|
||||
return super().金额(结果)
|
||||
|
||||
执行器 = 模型执行器(
|
||||
运行,
|
||||
预算,
|
||||
原文,
|
||||
证据服务(库),
|
||||
策略,
|
||||
直接宿主(HTTP传输(合成HTTP(回应), "受控存储", str(凭据), "responses")),
|
||||
会话计价(),
|
||||
)
|
||||
会话 = 角色会话(执行器, 上下文, 请求, 阶段="执行", 会话授权ID=会话授权, 工具登记=(定义,))
|
||||
if 宿主类型 in {"resume", "reconcile"}:
|
||||
if 宿主类型 == "resume":
|
||||
首回合 = asyncio.run(会话.模型回合())
|
||||
assert 首回合.状态 == "tool_calls"
|
||||
asyncio.run(会话.工具回合(首回合.工具调用[0]))
|
||||
运行.控制任务(上下文.领取.任务ID, "author", 任务状态.运行中, "暂停", 命令ID="pause")
|
||||
原状态 = 任务状态.已暂停
|
||||
else:
|
||||
with pytest.raises(预算不足):
|
||||
asyncio.run(会话.模型回合())
|
||||
运行.失败步骤(上下文.领取, "unknown_cost")
|
||||
with 库.连接() as 连:
|
||||
证据ID = str(
|
||||
连.execute(
|
||||
"SELECT evidence_id FROM muse_runtime_evidence WHERE kind='model_response'"
|
||||
).fetchone()[0]
|
||||
)
|
||||
预算.结算("call-1", Decimal("0.125"), 回执ID="session-response-1")
|
||||
运行.对账调用(
|
||||
上下文.领取.任务ID,
|
||||
上下文.领取.尝试ID,
|
||||
"author",
|
||||
已保存输出={"evidence_id": 证据ID},
|
||||
对账回执="priced-response-1",
|
||||
)
|
||||
assert 运行.读取任务(上下文.领取.任务ID).步骤[0]["state"] == "pending"
|
||||
原状态 = 任务状态.待对账
|
||||
运行.控制任务(上下文.领取.任务ID, "author", 原状态, "恢复", 命令ID="resume")
|
||||
新领取 = 运行.领取步骤("resumed-worker", ["model-probe"])
|
||||
assert 新领取 is not None and 新领取.尝试ID != 上下文.领取.尝试ID
|
||||
上下文 = 执行上下文(新领取, 运行.读取任务(新领取.任务ID))
|
||||
会话 = 角色会话(执行器, 上下文, 请求, 阶段="执行", 会话授权ID=会话授权, 工具登记=(定义,))
|
||||
if 宿主类型 == "direct":
|
||||
宿主 = 直接角色循环()
|
||||
else:
|
||||
from muse.基础设施.宿主.Pi import Pi宿主
|
||||
|
||||
# 显式配置缺失是失败,不能将替身或跳过计作Pi整链通过。
|
||||
宿主 = Pi宿主(
|
||||
Path(os.environ["MUSE_PI_NODE"]), Path(os.environ["MUSE_PI_PACKAGE"]), "0.84.4"
|
||||
)
|
||||
交付 = asyncio.run(会话.执行(宿主))
|
||||
assert 交付.内容 == {"text": "结构已读取,规划待作者确认"}
|
||||
assert len(收到) == 2 and 读取次数 == [上下文.领取.任务ID]
|
||||
assert 预算.读取("call-1").实际金额 == 预算.读取("call-1:round:2").实际金额 == Decimal("0.125")
|
||||
with 库.连接() as 连:
|
||||
assert 连.execute("SELECT count(*) FROM muse_runtime_evidence").fetchone()[0] == 5
|
||||
assert (
|
||||
连.execute(
|
||||
"SELECT count(*) FROM muse_raw_authorization WHERE parent_authorization_id=%s",
|
||||
(会话授权,),
|
||||
).fetchone()[0]
|
||||
== 3
|
||||
)
|
||||
行 = 连.execute(
|
||||
"SELECT content FROM muse_runtime_evidence WHERE kind='tool_result'"
|
||||
).fetchone()
|
||||
assert json.loads(行[0])["result"]["内容"]["schema_id"] == "work_core"
|
||||
运行.完成步骤(上下文.领取, 步骤结果({"evidence_id": 交付.证据回执["evidence_id"]}))
|
||||
assert 运行.读取任务(上下文.领取.任务ID).状态 == 任务状态.已完成
|
||||
|
||||
|
||||
@pytest.mark.parametrize(
|
||||
"场景", ["成功", "结构错误", "费用未知"], ids=["valid", "invalid-output", "unknown-cost"]
|
||||
)
|
||||
def test_受控模型流同时保存终态证据费用且拒绝重发__a62001(模型环境, tmp_path, 场景, 合成HTTP):
|
||||
库, 运行, 策略, 请求, 原文, 授权, 预算, 上下文 = 模型环境
|
||||
外发 = []
|
||||
|
||||
def 上游(req):
|
||||
外发.append(req)
|
||||
with 库.连接() as 连:
|
||||
# 上游收到请求时,尝试状态和预算已经在同一PG提交中登记。
|
||||
assert (
|
||||
连.execute(
|
||||
"SELECT call_state FROM muse_attempt WHERE attempt_id=%s", (上下文.领取.尝试ID,)
|
||||
).fetchone()[0]
|
||||
== "sent"
|
||||
)
|
||||
assert (
|
||||
连.execute(
|
||||
"SELECT state FROM muse_budget_reservation WHERE call_id='call-1'"
|
||||
).fetchone()[0]
|
||||
== "in_flight"
|
||||
)
|
||||
response = {
|
||||
"id": "response-1",
|
||||
"model": "claude-opus-4-8",
|
||||
"status": "completed",
|
||||
"output": [
|
||||
{
|
||||
"type": "message",
|
||||
"content": [
|
||||
{
|
||||
"type": "output_text",
|
||||
"text": '{"text":"合成内容"}' if 场景 != "结构错误" else "不是JSON",
|
||||
}
|
||||
],
|
||||
}
|
||||
],
|
||||
}
|
||||
if 场景 != "费用未知":
|
||||
response["usage"] = {"input_tokens": 10, "output_tokens": 12}
|
||||
payload = json.dumps(
|
||||
{"type": "response.completed", "response": response}, ensure_ascii=False
|
||||
)
|
||||
return ("data: " + payload + "\n\n").encode()
|
||||
|
||||
口令 = tmp_path / "synthetic-key"
|
||||
口令.write_text("synthetic-credential")
|
||||
口令.chmod(0o600)
|
||||
传输 = HTTP传输(合成HTTP(上游), "受控存储", str(口令), "responses")
|
||||
执行器 = 模型执行器(运行, 预算, 原文, 证据服务(库), 策略, 直接宿主(传输), 合成计价())
|
||||
|
||||
def 执行():
|
||||
return asyncio.run(执行器.执行(上下文, 请求, 阶段="生成", 原文授权ID=授权))
|
||||
|
||||
if 场景 == "成功":
|
||||
交付 = 执行()
|
||||
assert 交付.内容 == {"text": "合成内容"} and 交付.证据回执["retention"] == "full"
|
||||
with pytest.raises(预算状态冲突):
|
||||
执行()
|
||||
运行.完成步骤(上下文.领取, 步骤结果({"evidence_id": 交付.证据回执["evidence_id"]}))
|
||||
assert 运行.读取任务(上下文.领取.任务ID).状态 == 任务状态.已完成
|
||||
else:
|
||||
with pytest.raises(模型协议错误 if 场景 == "结构错误" else 预算不足):
|
||||
执行()
|
||||
运行.失败步骤(上下文.领取, "synthetic_failure")
|
||||
assert 运行.读取任务(上下文.领取.任务ID).状态 == (
|
||||
任务状态.已失败 if 场景 == "结构错误" else 任务状态.待对账
|
||||
)
|
||||
assert len(外发) == 1
|
||||
with 库.连接() as 连:
|
||||
行 = 连.execute(
|
||||
"SELECT outcome,content FROM muse_runtime_evidence WHERE kind='model_response'"
|
||||
).fetchone()
|
||||
assert 行 is not None and 行[0] == ("failed" if 场景 == "结构错误" else "completed")
|
||||
assert 行[1] is not None
|
||||
assert 预算.读取("call-1").实际金额 == (None if 场景 == "费用未知" else Decimal("0.125"))
|
||||
|
||||
|
||||
def test_调用授权不符在预算和模型前拒绝__a62002(模型环境):
|
||||
from dataclasses import replace
|
||||
|
||||
库, 运行, 策略, 请求, 原文, 授权, 预算, 上下文 = 模型环境
|
||||
|
||||
class 不应调用:
|
||||
def 准备(self, *args, **kwargs):
|
||||
raise AssertionError("不应到达模型")
|
||||
|
||||
执行器 = 模型执行器(运行, 预算, 原文, 证据服务(库), 策略, 不应调用(), 合成计价())
|
||||
with pytest.raises(原文错误):
|
||||
asyncio.run(
|
||||
执行器.执行(
|
||||
上下文, replace(请求, 用户输入="被替换的输入"), 阶段="生成", 原文授权ID=授权
|
||||
)
|
||||
)
|
||||
with 库.连接() as 连:
|
||||
assert 连.execute("SELECT count(*) FROM muse_budget_reservation").fetchone()[0] == 0
|
||||
assert 连.execute("SELECT count(*) FROM muse_runtime_evidence").fetchone()[0] == 0
|
||||
|
||||
|
||||
def test_任务资源与当前构建不符时拒绝调用__a62004(模型环境):
|
||||
库, 运行, 策略, 请求, 原文, 授权, 预算, 上下文 = 模型环境
|
||||
# 已创建任务保留原构建身份,模拟执行器换装了另一个资源构建。
|
||||
策略.资源发布身份 = "0" * 64
|
||||
|
||||
class 不应准备:
|
||||
def 准备(self, *args, **kwargs):
|
||||
raise AssertionError("资源漂移必须在宿主准备前拒绝")
|
||||
|
||||
执行器 = 模型执行器(运行, 预算, 原文, 证据服务(库), 策略, 不应准备(), 合成计价())
|
||||
with pytest.raises(模型协议错误, match="资源"):
|
||||
asyncio.run(执行器.执行(上下文, 请求, 阶段="生成", 原文授权ID=授权))
|
||||
with 库.连接() as 连:
|
||||
assert 连.execute("SELECT count(*) FROM muse_budget_reservation").fetchone()[0] == 0
|
||||
|
||||
|
||||
@pytest.mark.parametrize("场景", ["missing-credential", "unsupported-protocol", "invalid-endpoint"])
|
||||
def test_发送前准备失败不记已调用或占未知费用__a62003(模型环境, tmp_path, 场景, 合成HTTP):
|
||||
from muse.共享.错误 import Muse错误
|
||||
|
||||
库, 运行, 策略, 请求, 原文, 授权, 预算, 上下文 = 模型环境
|
||||
外发 = []
|
||||
地址 = 合成HTTP(lambda value: 外发.append(value))
|
||||
凭据 = tmp_path / "preflight-credential"
|
||||
if 场景 != "missing-credential":
|
||||
凭据.write_text("synthetic-credential")
|
||||
宿主 = 直接宿主(
|
||||
HTTP传输(
|
||||
"not-an-endpoint" if 场景 == "invalid-endpoint" else 地址,
|
||||
"受控存储",
|
||||
str(凭据),
|
||||
"unknown" if 场景 == "unsupported-protocol" else "responses",
|
||||
)
|
||||
)
|
||||
执行器 = 模型执行器(运行, 预算, 原文, 证据服务(库), 策略, 宿主, 合成计价())
|
||||
with pytest.raises(Muse错误):
|
||||
asyncio.run(执行器.执行(上下文, 请求, 阶段="生成", 原文授权ID=授权))
|
||||
assert not 外发
|
||||
with 库.连接() as 连:
|
||||
assert 连.execute("SELECT count(*) FROM muse_budget_reservation").fetchone()[0] == 0
|
||||
assert 连.execute("SELECT count(*) FROM muse_runtime_evidence").fetchone()[0] == 0
|
||||
assert (
|
||||
连.execute(
|
||||
"SELECT call_state FROM muse_attempt WHERE attempt_id=%s", (上下文.领取.尝试ID,)
|
||||
).fetchone()[0]
|
||||
!= "sent"
|
||||
)
|
||||
运行.失败步骤(上下文.领取, "configuration_invalid")
|
||||
assert 运行.读取任务(上下文.领取.任务ID).状态 == 任务状态.已失败
|
||||
331
tests/集成/test_预算预留与结算.py
Normal file
331
tests/集成/test_预算预留与结算.py
Normal file
@ -0,0 +1,331 @@
|
||||
"""真实 PG 的预算原子预留、成本状态与配置版本;不调用外部模型。"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import dataclasses
|
||||
import time
|
||||
import uuid
|
||||
from concurrent.futures import ThreadPoolExecutor
|
||||
from datetime import UTC, datetime, timedelta
|
||||
from decimal import Decimal
|
||||
from pathlib import Path
|
||||
from threading import Barrier
|
||||
|
||||
import psycopg
|
||||
import pytest
|
||||
from psycopg import sql
|
||||
|
||||
from muse.任务运行.接口 import 任务服务, 任务请求, 步骤处理器, 步骤结果, 步骤计划
|
||||
from muse.任务运行.模型 import 内容哈希
|
||||
from muse.任务运行.配置版本 import (
|
||||
凭据引用,
|
||||
提供方配置,
|
||||
运行配置内容,
|
||||
配置版本管理,
|
||||
配置版本错误,
|
||||
配置验证证据,
|
||||
)
|
||||
from muse.任务运行.预算管理 import (
|
||||
任务预算计划,
|
||||
角色预算,
|
||||
预算不足,
|
||||
预算状态冲突,
|
||||
预算管理,
|
||||
额度策略,
|
||||
)
|
||||
from muse.共享.调用身份 import 内容用途, 用途
|
||||
from muse.基础设施.数据库.迁移 import 执行迁移
|
||||
from muse.基础设施.数据库.连接 import 数据库工厂
|
||||
from muse.编排.接口 import 流程定义, 流程服务, 流程登记
|
||||
from muse.配置 import 数据库引用
|
||||
|
||||
pytestmark = pytest.mark.数据库
|
||||
根 = Path(__file__).resolve().parents[2]
|
||||
|
||||
|
||||
@pytest.fixture
|
||||
def 预算环境(隔离数据库URL: str, monkeypatch: pytest.MonkeyPatch, tmp_path: Path):
|
||||
库名 = f"muse_budget_{uuid.uuid4().hex[:12]}"
|
||||
with psycopg.connect(隔离数据库URL, autocommit=True) as 连:
|
||||
连.execute((根 / "数据库/初始化/用途角色.sql").read_text())
|
||||
连.execute(
|
||||
sql.SQL("CREATE DATABASE {} OWNER muse_maint TEMPLATE template0").format(
|
||||
sql.Identifier(库名)
|
||||
)
|
||||
)
|
||||
工厂 = {}
|
||||
try:
|
||||
for 声明, 角色 in [
|
||||
(用途.维护, "muse_maint"),
|
||||
(用途.生产, "muse_app"),
|
||||
(用途.评测, "muse_eval"),
|
||||
]:
|
||||
参数 = psycopg.conninfo.conninfo_to_dict(隔离数据库URL)
|
||||
参数.update(dbname=库名, user=角色)
|
||||
环境名 = f"MUSE_BUDGET_{声明.name}_URL"
|
||||
monkeypatch.setenv(环境名, psycopg.conninfo.make_conninfo(**参数))
|
||||
工厂[声明] = 数据库工厂(数据库引用("环境变量", 环境名), 声明)
|
||||
for 名称 in [
|
||||
"V0001__共享标识与版本.sql",
|
||||
"V0004__任务运行与证据.sql",
|
||||
"V0016__流程定义与版本.sql",
|
||||
]:
|
||||
(tmp_path / 名称).write_bytes((根 / "数据库/迁移" / 名称).read_bytes())
|
||||
with 工厂[用途.维护].连接() as 连:
|
||||
执行迁移(连, tmp_path)
|
||||
yield 工厂
|
||||
finally:
|
||||
with psycopg.connect(隔离数据库URL, autocommit=True) as 连:
|
||||
连.execute(sql.SQL("DROP DATABASE {} WITH (FORCE)").format(sql.Identifier(库名)))
|
||||
|
||||
|
||||
def 新任务(工厂: 数据库工厂):
|
||||
登记 = 流程登记()
|
||||
登记.登记处理器(步骤处理器("call", "1", lambda _: 步骤结果({}), "v1", "v1"))
|
||||
登记.登记类型("budget-test", 必需保护=())
|
||||
运行 = 任务服务(工厂, 登记)
|
||||
编排 = 流程服务(运行, 登记)
|
||||
编排.发布(流程定义("budget-test", "1", (步骤计划("call", "call", "1"),)))
|
||||
请求 = 任务请求(
|
||||
"budget-test",
|
||||
uuid.uuid4().hex,
|
||||
"author",
|
||||
工厂.用途,
|
||||
内容用途.生成,
|
||||
{},
|
||||
"policy-1",
|
||||
"release-1",
|
||||
{
|
||||
"source_scope": {},
|
||||
"schema_versions": {},
|
||||
"authorization": "grant",
|
||||
"budget": {"approved": True},
|
||||
"stop_conditions": ["cancelled"],
|
||||
},
|
||||
)
|
||||
身份 = 编排.发起(请求, "budget-test", "1")
|
||||
领取 = 运行.领取步骤("worker", ["call"])
|
||||
assert 领取 is not None and 领取.任务ID == 身份
|
||||
return 身份, 领取
|
||||
|
||||
|
||||
def 建预算(工厂: 数据库工厂, 任务ID: str, *, 窗口金额="24", 单次="6", 次数=3, 窗口次数=6000):
|
||||
管理 = 预算管理(工厂, "quota")
|
||||
管理.登记策略(额度策略("quota", "1", Decimal(窗口金额), 窗口次数))
|
||||
管理.登记任务预算(
|
||||
任务ID,
|
||||
任务预算计划(
|
||||
Decimal(单次) * 次数,
|
||||
(角色预算("writer", 次数, 次数, Decimal(单次)),),
|
||||
"approval",
|
||||
datetime.now(UTC) + timedelta(minutes=10),
|
||||
),
|
||||
)
|
||||
return 管理
|
||||
|
||||
|
||||
def test_并发预留共享窗口且不超额__a61001(预算环境) -> None:
|
||||
A, 领取A = 新任务(预算环境[用途.生产])
|
||||
B, 领取B = 新任务(预算环境[用途.生产])
|
||||
预算A = 建预算(预算环境[用途.生产], A, 窗口金额="10")
|
||||
预算B = 建预算(预算环境[用途.生产], B, 窗口金额="10")
|
||||
栅栏 = Barrier(2)
|
||||
|
||||
def 预留(输入):
|
||||
管理, 领取 = 输入
|
||||
栅栏.wait()
|
||||
try:
|
||||
return 管理.预留(领取, 领取.任务ID, "writer")
|
||||
except 预算不足:
|
||||
return None
|
||||
|
||||
with ThreadPoolExecutor(max_workers=2) as 线程:
|
||||
结果 = list(线程.map(预留, [(预算A, 领取A), (预算B, 领取B)]))
|
||||
assert sum(r is not None for r in 结果) == 1
|
||||
assert 预算A.窗口余额()["在途预留"] == Decimal("6")
|
||||
assert 预算A.窗口余额()["可用金额"] == Decimal("4")
|
||||
|
||||
|
||||
def test_重复预留发送和结算不重复计费__a61002(预算环境) -> None:
|
||||
身份, 领取 = 新任务(预算环境[用途.生产])
|
||||
预算 = 建预算(预算环境[用途.生产], 身份)
|
||||
初次 = 预算.预留(领取, "call-1", "writer")
|
||||
assert 预算.预留(领取, "call-1", "writer") == 初次
|
||||
assert 预算.标记已发送(领取, "call-1", 最长秒=30).允许外发
|
||||
assert not 预算.标记已发送(领取, "call-1", 最长秒=30).允许外发
|
||||
结果 = 预算.结算("call-1", Decimal("0.125"), 回执ID="failed-call-receipt")
|
||||
assert 预算.结算("call-1", Decimal("0.125"), 回执ID="failed-call-receipt") == 结果
|
||||
with pytest.raises(预算状态冲突):
|
||||
预算.结算("call-1", Decimal("0.2"), 回执ID="changed")
|
||||
assert 预算.窗口余额()["已知成本"] == Decimal("0.125")
|
||||
assert 预算.窗口余额()["已占次数"] == 1
|
||||
|
||||
|
||||
def test_未知成本跨窗口仍阻止后续调用直至对账__a61003(预算环境) -> None:
|
||||
身份, 领取 = 新任务(预算环境[用途.生产])
|
||||
预算 = 建预算(预算环境[用途.生产], 身份)
|
||||
预算.预留(领取, "unknown", "writer")
|
||||
预算.标记已发送(领取, "unknown", 最长秒=30)
|
||||
未知 = 预算.结算("unknown", None, 回执ID="missing-cost")
|
||||
assert 未知.实际金额 is None and 未知.状态 == "unknown"
|
||||
with 预算环境[用途.维护].连接() as 连:
|
||||
连.execute(
|
||||
"UPDATE public.muse_budget_reservation SET window_start=window_start-interval '1 "
|
||||
"day',window_end=window_end-interval '1 day' WHERE call_id='unknown'"
|
||||
)
|
||||
assert 预算.窗口余额()["未知调用数"] == 1
|
||||
with pytest.raises(预算不足, match="未知成本"):
|
||||
预算.预留(领取, "next", "writer")
|
||||
预算.结算("unknown", Decimal("0.2"), 回执ID="resolved-receipt")
|
||||
assert 预算.预留(领取, "next", "writer").状态 == "reserved"
|
||||
|
||||
|
||||
@pytest.mark.parametrize("已发出", [False, True], ids=["before-send", "after-send"])
|
||||
def test_取消保留已发送调用的未知成本__a61004(预算环境, 已发出: bool) -> None:
|
||||
身份, 领取 = 新任务(预算环境[用途.生产])
|
||||
预算 = 建预算(预算环境[用途.生产], 身份)
|
||||
预算.预留(领取, "cancel", "writer")
|
||||
if 已发出:
|
||||
预算.标记已发送(领取, "cancel", 最长秒=30)
|
||||
预算.关闭任务预算(身份)
|
||||
assert 预算.读取("cancel").状态 == ("unknown" if 已发出 else "released")
|
||||
with pytest.raises(预算不足):
|
||||
预算.预留(领取, "blocked", "writer")
|
||||
|
||||
|
||||
@pytest.mark.parametrize("已发出", [False, True], ids=["reserved", "in-flight"])
|
||||
def test_过期只释放尚未发送的预留__a61005(预算环境, 已发出: bool) -> None:
|
||||
身份, 领取 = 新任务(预算环境[用途.生产])
|
||||
预算 = 建预算(预算环境[用途.生产], 身份)
|
||||
预算.预留(领取, "expire", "writer", 有效秒=0.08 if not 已发出 else 30)
|
||||
if 已发出:
|
||||
预算.标记已发送(领取, "expire", 最长秒=0.08)
|
||||
time.sleep(0.12)
|
||||
预算.回收过期()
|
||||
assert 预算.读取("expire").状态 == ("unknown" if 已发出 else "released")
|
||||
if not 已发出:
|
||||
assert 预算.窗口余额()["在途预留"] == 0
|
||||
assert 预算.预留(领取, "fresh", "writer").状态 == "reserved"
|
||||
|
||||
|
||||
def test_超单次成本先保存事实并关闭任务预算__a61006(预算环境) -> None:
|
||||
身份, 领取 = 新任务(预算环境[用途.生产])
|
||||
预算 = 建预算(预算环境[用途.生产], 身份)
|
||||
预算.预留(领取, "over", "writer")
|
||||
预算.标记已发送(领取, "over", 最长秒=30)
|
||||
已记 = 预算.结算("over", Decimal("7"), 回执ID="actual-over")
|
||||
assert 已记.超预算 and 已记.实际金额 == Decimal("7")
|
||||
assert 预算.窗口余额()["已知成本"] == Decimal("7")
|
||||
with pytest.raises(预算不足):
|
||||
预算.预留(领取, "next", "writer")
|
||||
|
||||
|
||||
def test_结算失败回滚仍保留在途标记__a61007(预算环境) -> None:
|
||||
身份, 领取 = 新任务(预算环境[用途.生产])
|
||||
预算 = 建预算(预算环境[用途.生产], 身份)
|
||||
预算.预留(领取, "atomic", "writer")
|
||||
预算.标记已发送(领取, "atomic", 最长秒=30)
|
||||
with 预算环境[用途.维护].连接() as 连:
|
||||
连.execute(
|
||||
"CREATE FUNCTION public.reject_settlement() RETURNS trigger LANGUAGE plpgsql AS $$ "
|
||||
"BEGIN IF NEW.state='settled' THEN RAISE EXCEPTION 'synthetic failure'; END IF; "
|
||||
"RETURN NEW; END $$"
|
||||
)
|
||||
连.execute(
|
||||
"CREATE TRIGGER reject_settlement BEFORE UPDATE ON public.muse_budget_reservation "
|
||||
"FOR EACH ROW EXECUTE FUNCTION public.reject_settlement()"
|
||||
)
|
||||
with pytest.raises(psycopg.errors.RaiseException):
|
||||
预算.结算("atomic", Decimal("1"), 回执ID="receipt")
|
||||
assert 预算.读取("atomic").状态 == "in_flight"
|
||||
assert 预算.读取("atomic").实际金额 is None
|
||||
with 预算环境[用途.维护].连接() as 连:
|
||||
连.execute("DROP TRIGGER reject_settlement ON public.muse_budget_reservation")
|
||||
assert 预算.结算("atomic", Decimal("1"), 回执ID="receipt").状态 == "settled"
|
||||
|
||||
|
||||
@pytest.mark.parametrize("模式", ["window", "task"], ids=["window", "task"])
|
||||
def test_窗口次数和任务计划分别限制调用__a61008(预算环境, 模式: str) -> None:
|
||||
身份, 领取 = 新任务(预算环境[用途.生产])
|
||||
预算 = 建预算(
|
||||
预算环境[用途.生产],
|
||||
身份,
|
||||
窗口次数=1 if 模式 == "window" else 6000,
|
||||
次数=3 if 模式 == "window" else 1,
|
||||
)
|
||||
预算.预留(领取, "first", "writer")
|
||||
预算.标记已发送(领取, "first", 最长秒=30)
|
||||
预算.结算("first", Decimal("0"), 回执ID="free-but-called")
|
||||
with pytest.raises(预算不足, match="次数"):
|
||||
预算.预留(领取, "second", "writer")
|
||||
|
||||
|
||||
class 合成配置验证器:
|
||||
身份 = "synthetic-contract-check-v1"
|
||||
|
||||
def 验证(self, 内容: 运行配置内容, 执行用途: 用途) -> 配置验证证据:
|
||||
assert 内容.角色配置["writer"]["model"] == "model"
|
||||
return 配置验证证据(
|
||||
内容哈希(内容.冻结()),
|
||||
内容.角色策略版本,
|
||||
内容.资源发布身份,
|
||||
执行用途,
|
||||
("synthetic-protocol-result",),
|
||||
"offline_contract",
|
||||
)
|
||||
|
||||
|
||||
def 配置内容(版本="1"):
|
||||
return 运行配置内容(
|
||||
"direct",
|
||||
版本,
|
||||
"policy-1",
|
||||
"release-1",
|
||||
"quota-1",
|
||||
{"writer": {"provider": "provider", "model": "model", "thinking": "high"}},
|
||||
(凭据引用("api", "环境变量", "MUSE_API_KEY"),),
|
||||
(提供方配置("provider", "responses", "https://provider.invalid/v1/responses", "api"),),
|
||||
"price-1",
|
||||
)
|
||||
|
||||
|
||||
def test_配置验证启用与任务冻结版本互不覆盖__a61009(预算环境) -> None:
|
||||
管理 = 配置版本管理(预算环境[用途.评测], 合成配置验证器())
|
||||
A, _ = 新任务(预算环境[用途.评测])
|
||||
原 = 管理.保存草案("config", "1", 配置内容())
|
||||
with pytest.raises(配置版本错误):
|
||||
管理.冻结到任务(A, "config")
|
||||
with pytest.raises(配置版本错误):
|
||||
管理.保存草案("config", "1", 配置内容("changed"))
|
||||
R1 = 管理.验证版本("config", "1")
|
||||
assert 管理.启用("config", "1", 验证回执=R1, 批准引用="approve-1", 预期代次=0) == 1
|
||||
assert 管理.启用("config", "1", 验证回执=R1, 批准引用="approve-1", 预期代次=0) == 1
|
||||
冻结 = 管理.冻结到任务(A, "config")
|
||||
管理.保存草案("config", "2", 配置内容("2"))
|
||||
R2 = 管理.验证版本("config", "2")
|
||||
with pytest.raises(配置版本错误):
|
||||
管理.启用("config", "2", 验证回执=R1, 批准引用="approve-2", 预期代次=1)
|
||||
assert 管理.启用("config", "2", 验证回执=R2, 批准引用="approve-2", 预期代次=1) == 2
|
||||
assert 管理.冻结到任务(A, "config") == 冻结 == 原
|
||||
B, _ = 新任务(预算环境[用途.评测])
|
||||
assert 管理.冻结到任务(B, "config").版本 == "2"
|
||||
管理.停用("config", 预期代次=2)
|
||||
C, _ = 新任务(预算环境[用途.评测])
|
||||
with pytest.raises(配置版本错误):
|
||||
管理.冻结到任务(C, "config")
|
||||
assert 管理.冻结到任务(A, "config").版本 == "1"
|
||||
|
||||
|
||||
def test_配置不接受陈旧验证或离线证据启用生产__a6100a(预算环境) -> None:
|
||||
class 陈旧验证器(合成配置验证器):
|
||||
def 验证(self, 内容, 执行用途):
|
||||
return dataclasses.replace(super().验证(内容, 执行用途), 配置哈希="old-hash")
|
||||
|
||||
管理 = 配置版本管理(预算环境[用途.评测], 陈旧验证器())
|
||||
管理.保存草案("config", "1", 配置内容())
|
||||
with pytest.raises(配置版本错误, match="绑定"):
|
||||
管理.验证版本("config", "1")
|
||||
生产 = 配置版本管理(预算环境[用途.生产], 合成配置验证器())
|
||||
生产.保存草案("config", "1", 配置内容())
|
||||
with pytest.raises(配置版本错误, match="离线"):
|
||||
生产.验证版本("config", "1")
|
||||
102
工具/环境预检.py
Normal file
102
工具/环境预检.py
Normal file
@ -0,0 +1,102 @@
|
||||
"""核对已安装资源、显式配置和本地宿主;不会发起模型调用。"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import argparse
|
||||
import json
|
||||
import os
|
||||
import subprocess
|
||||
import tempfile
|
||||
from pathlib import Path
|
||||
|
||||
from muse.共享.错误 import Muse错误, 配置错误
|
||||
from muse.启动 import 构建
|
||||
from muse.资源加载 import 核对资源, 读取能力目录
|
||||
from muse.配置 import 读取配置
|
||||
|
||||
|
||||
def 检查Pi(node: Path, cli: Path, 预期版本: str) -> dict:
|
||||
if not node.is_file() or not cli.is_file():
|
||||
raise 配置错误("Node或Pi入口不存在;需要明确的本地可执行路径")
|
||||
结果 = subprocess.run(
|
||||
[str(node), "--version"], capture_output=True, text=True, timeout=10, check=True
|
||||
)
|
||||
node_version = 结果.stdout.strip().removeprefix("v")
|
||||
parts = tuple(int(v) for v in node_version.split("."))
|
||||
if parts < (22, 19, 0):
|
||||
raise 配置错误("当前Pi需要Node至少22.19.0;不能使用shell默认旧版本")
|
||||
with tempfile.TemporaryDirectory(prefix="muse-pi-preflight-") as 临时:
|
||||
环境 = {
|
||||
**os.environ,
|
||||
"PI_CODING_AGENT_DIR": 临时,
|
||||
"PI_OFFLINE": "1",
|
||||
"PI_TELEMETRY": "0",
|
||||
}
|
||||
结果 = subprocess.run(
|
||||
[str(node), str(cli), "--version"],
|
||||
cwd=临时,
|
||||
env=环境,
|
||||
capture_output=True,
|
||||
text=True,
|
||||
timeout=15,
|
||||
check=True,
|
||||
)
|
||||
实际 = 结果.stdout.strip()
|
||||
if 实际 != 预期版本:
|
||||
raise 配置错误("Pi实际版本与显式要求不符")
|
||||
return {"node_version": node_version, "pi_version": 实际, "state": "version_verified"}
|
||||
|
||||
|
||||
def 主() -> int:
|
||||
解析 = argparse.ArgumentParser(description="检查固定资源和明确选择的运行依赖,不调用模型")
|
||||
解析.add_argument("--配置", type=Path)
|
||||
解析.add_argument("--数据库", action="store_true", help="明确执行所给配置的数据库连接检查")
|
||||
解析.add_argument("--node", type=Path)
|
||||
解析.add_argument("--pi", type=Path)
|
||||
解析.add_argument("--pi-version")
|
||||
参数 = 解析.parse_args()
|
||||
if 参数.数据库 and 参数.配置 is None:
|
||||
解析.error("数据库检查需要明确配置")
|
||||
if any((参数.node, 参数.pi, 参数.pi_version)) and not all(
|
||||
(参数.node, 参数.pi, 参数.pi_version)
|
||||
):
|
||||
解析.error("Pi检查必须同时给出node、pi和pi-version")
|
||||
try:
|
||||
清单 = 核对资源()
|
||||
能力 = 读取能力目录()
|
||||
报告 = {
|
||||
"resource_build_id": 清单["构建身份"],
|
||||
"resource_label": 清单["发布身份"],
|
||||
"package_version": 清单["代码版本"],
|
||||
"resources_verified": len(清单["资源"]),
|
||||
"roles": sorted(能力["role"]),
|
||||
"operations": sorted(能力["operation"]),
|
||||
"database": "not_checked",
|
||||
"host": {"state": "not_checked"},
|
||||
"model_calls": 0,
|
||||
}
|
||||
if 参数.配置:
|
||||
配置 = 读取配置(参数.配置)
|
||||
if 配置.资源发布身份 != 清单["构建身份"]:
|
||||
raise 配置错误("配置必须固定当前资源构建身份;先无配置运行预检取得身份")
|
||||
if 参数.数据库:
|
||||
构建(配置).要求数据库().检查连接()
|
||||
报告["database"] = "connection_verified"
|
||||
if 参数.node:
|
||||
报告["host"] = 检查Pi(参数.node.resolve(), 参数.pi.resolve(), 参数.pi_version)
|
||||
print(json.dumps(报告, ensure_ascii=False, indent=2))
|
||||
return 0
|
||||
except Muse错误 as exc:
|
||||
print(json.dumps(exc.呈现(), ensure_ascii=False))
|
||||
except (OSError, ValueError, KeyError, subprocess.SubprocessError):
|
||||
print(
|
||||
json.dumps(
|
||||
{"code": "PREFLIGHT_FAILED", "message": "资源或宿主检查失败,未验证项目不计通过"},
|
||||
ensure_ascii=False,
|
||||
)
|
||||
)
|
||||
return 1
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
raise SystemExit(主())
|
||||
25
配置/应用.example.toml
Normal file
25
配置/应用.example.toml
Normal file
@ -0,0 +1,25 @@
|
||||
# 复制到私人配置位置;连接值保存在受控文件,不写入此模板。
|
||||
["数据库"]
|
||||
"取值方式" = "受控存储"
|
||||
# 相对地址以配置文件所在目录为起点。
|
||||
"位置" = "数据库连接.txt"
|
||||
|
||||
["资源"]
|
||||
# 从 工具/环境预检.py 的 resource_build_id 取得,不能只写发布标签。
|
||||
"发布身份" = "<固定资源构建身份>"
|
||||
|
||||
["运行"]
|
||||
# 生产/评测/维护各用独立角色连接;显式迁移必须 maintenance。
|
||||
"用途" = "production"
|
||||
|
||||
[HTTP]
|
||||
"作者ID" = "author-local"
|
||||
"口令文件" = "工作台口令.txt"
|
||||
"监听地址" = "127.0.0.1"
|
||||
"端口" = 8080
|
||||
"公开地址" = "http://127.0.0.1:8080"
|
||||
"允许来源" = ["http://127.0.0.1:8080"]
|
||||
|
||||
["文件"]
|
||||
"原文暂存" = "证据暂存"
|
||||
"探索草稿" = "私人探索"
|
||||
29
配置/提供方.example.toml
Normal file
29
配置/提供方.example.toml
Normal file
@ -0,0 +1,29 @@
|
||||
# 本文件是运行配置草案,保存不等于验证或启用。不填写凭据值。
|
||||
"宿主" = "direct"
|
||||
"宿主版本" = "1"
|
||||
"角色策略版本" = "role-policy-r2-v1"
|
||||
"资源发布身份" = "<环境预检返回的resource_build_id>"
|
||||
"预算策略引用" = "<已登记且不可变的预算账户身份>"
|
||||
"计价版本" = "<后端已登记计价实现的版本>"
|
||||
# Pi模式同时指定固定安装,版本改为实际批准版本:
|
||||
# "宿主" = "pi"
|
||||
# "宿主版本" = "0.84.4"
|
||||
# "Node路径" = "/absolute/path/to/node"
|
||||
# "Pi包目录" = "/absolute/path/to/pi-coding-agent"
|
||||
|
||||
["角色配置".writer]
|
||||
provider = "primary"
|
||||
model = "claude-opus-4-8[1M]"
|
||||
thinking = "high"
|
||||
tools = []
|
||||
|
||||
[["凭据"]]
|
||||
"名称" = "primary-key"
|
||||
"来源" = "受控存储"
|
||||
"位置" = "/absolute/private/path/provider-key.txt"
|
||||
|
||||
[["提供方"]]
|
||||
"身份" = "primary"
|
||||
"协议" = "responses"
|
||||
"地址" = "https://provider.invalid/v1/responses"
|
||||
"凭据名称" = "primary-key"
|
||||
21
配置/角色策略.yaml
Normal file
21
配置/角色策略.yaml
Normal file
@ -0,0 +1,21 @@
|
||||
version: role-policy-r2-v1
|
||||
aliases:
|
||||
blind_judge: judge
|
||||
semantic_detector: detector
|
||||
models:
|
||||
fixed:
|
||||
- claude-opus-4-8[1M]
|
||||
governed:
|
||||
- claude-opus-4-8[1M]
|
||||
- MiniMax-M3
|
||||
- MiniMax-M2.7
|
||||
- glm-5.2
|
||||
- deepseek-v4-flash
|
||||
actual_model_ids:
|
||||
claude-opus-4-8[1M]: ["claude-opus-4-8[1M]", "claude-opus-4-8"]
|
||||
roles:
|
||||
writer: {model_policy: fixed, readonly_tools: true}
|
||||
planner: {model_policy: fixed, readonly_tools: true}
|
||||
judge: {model_policy: fixed, readonly_tools: true}
|
||||
detector: {model_policy: governed, readonly_tools: true}
|
||||
extractor: {model_policy: governed, readonly_tools: true}
|
||||
Loading…
x
Reference in New Issue
Block a user