重构(测试): 删除85次合成调用的资格终态用例

去掉四条用假HTTP循环85次的慢用例,以及浏览器报告中依赖同一路径的effect参数。
资格门槛仍由契约与单元覆盖;生产效果标准不改。
This commit is contained in:
zizi 2026-09-18 11:21:41 +08:00
parent 3943a4cdf9
commit 0eba8e4f34
8 changed files with 11 additions and 764 deletions

View File

@ -5,7 +5,7 @@
1. **默认离线**:默认测试进程由 macOS sandbox 或 Linux seccomp 禁止外部连接,覆盖收集阶段、数据库驱动及其子进程;防护不可用时失败,不退回无防护执行。显式外部用例使用 `--外部环境` 与环境标记选择。 1. **默认离线**:默认测试进程由 macOS sandbox 或 Linux seccomp 禁止外部连接,覆盖收集阶段、数据库驱动及其子进程;防护不可用时失败,不退回无防护执行。显式外部用例使用 `--外部环境` 与环境标记选择。
2. **显式外部环境**:需要 PostgreSQL 的用例标记 `数据库` 并使用对应夹具;缺 `MUSE_TEST_DATABASE_URL` 时直接失败,不跳过或回退旧库。普通业务测试默认使用同一会话的共享已迁移测试库(例间在同一连接清空业务表并重设序列,保留模板与结构种子),不再为每条普通用例执行 `CREATE DATABASE ... TEMPLATE` 克隆与全量销毁;仅当用例自身验证目标为数据库创建、迁移升级、跨版本切换或模板隔离时,才使用独立克隆库。 2. **显式外部环境**:需要 PostgreSQL 的用例标记 `数据库` 并使用对应夹具;缺 `MUSE_TEST_DATABASE_URL` 时直接失败,不跳过或回退旧库。普通业务测试默认使用同一会话的共享已迁移测试库(例间在同一连接清空业务表并重设序列,保留模板与结构种子),不再为每条普通用例执行 `CREATE DATABASE ... TEMPLATE` 克隆与全量销毁;仅当用例自身验证目标为数据库创建、迁移升级、跨版本切换或模板隔离时,才使用独立克隆库。
超过三分钟的大 N 统计评测用例标记 `慢`,属于评测实验范畴,默认不包含在常规数据库验证中,只有用户显式请求实验时才运行。 超过三分钟的用例标记 `慢`,默认不进常规数据库验证;浏览器旅程等慢入口走对应显式目标,不另设合成调用资格终态套件。
实际Pi循环另标记`宿主`,通过`make 宿主测试`显式提供固定Node与Pi包路径;普通数据库测试剔除此标记。宿主验证使用合成提供方,不等于真实模型验证。真实外部模型另标记`真实模型`,通过`make 真实模型测试`显式提供地址与凭据文件;普通数据库测试剔除该标记。 实际Pi循环另标记`宿主`,通过`make 宿主测试`显式提供固定Node与Pi包路径;普通数据库测试剔除此标记。宿主验证使用合成提供方,不等于真实模型验证。真实外部模型另标记`真实模型`,通过`make 真实模型测试`显式提供地址与凭据文件;普通数据库测试剔除该标记。
3. **用例身份**:完整 `case_id` 与清单中的实际文件、符号及参数行绑定,不靠名称后缀匹配;`pytest --case TC-…` 选择用例,`--case 'TC-…[参数ID]'` 选择参数行。同符号承接多个 ID 时必须分别绑定无交叠参数。未知 ID、目标缺失与冒用均失败;Junit 留存完整 ID 和参数 ID。 3. **用例身份**:完整 `case_id` 与清单中的实际文件、符号及参数行绑定,不靠名称后缀匹配;`pytest --case TC-…` 选择用例,`--case 'TC-…[参数ID]'` 选择参数行。同符号承接多个 ID 时必须分别绑定无交叠参数。未知 ID、目标缺失与冒用均失败;Junit 留存完整 ID 和参数 ID。
4. **可判定性**:用例体内必须有 assert/raise/fail/skip/xfail;纯 print 用例被索引检查在执行前拦截。 4. **可判定性**:用例体内必须有 assert/raise/fail/skip/xfail;纯 print 用例被索引检查在执行前拦截。

View File

@ -70,7 +70,7 @@ SoT 按主题分域,不做跨主题的全局排序。可执行脚本与书面
## 4. 工作协议(硬约束) ## 4. 工作协议(硬约束)
1. **读后动手与渐进发现**:复杂任务先读对应 SoT;涉及角色时读角色合同。技能发现遵守本文件开头的唯一入口,按当前任务展开,不预载无关资料。 1. **读后动手与渐进发现**:复杂任务先读对应 SoT;涉及角色时读角色合同。技能发现遵守本文件开头的唯一入口,按当前任务展开,不预载无关资料。
2. **机械验证优先与完成=验证**:工程统一使用仓内解释器 `.venv`(uv 管理);日常唯一入口为 `make 快检 范围=<受影响路径>`(秒级,开发期);离线全套 `make 测试`(~70秒);仅改动存储/迁移/触发器时跑对应模块的库测试。**全量数据库测试高耗时高耗盘,严禁未经用户明确请求擅自执行**:`make 验收数据库` 与全量 `make 数据库测试` 必须由用户显式触发(需传入 `MUSE_RUN_FULL_DB=1` 确认);慢评测实验(85次模型调用的终态验证)默认不跑,由 `make 评测实验` 显式申请;浏览器层走 `make 浏览器测试`。日常以模块快检或模块局部数据库测试为完成证据,无证据严禁声称“完成/修复/通过”。 2. **机械验证优先与完成=验证**:工程统一使用仓内解释器 `.venv`(uv 管理);日常唯一入口为 `make 快检 范围=<受影响路径>`(秒级,开发期);离线全套 `make 测试`(~70秒);仅改动存储/迁移/触发器时跑对应模块的库测试。**全量数据库测试高耗时高耗盘,严禁未经用户明确请求擅自执行**:`make 验收数据库` 与全量 `make 数据库测试` 必须由用户显式触发(需传入 `MUSE_RUN_FULL_DB=1` 确认);浏览器层走 `make 浏览器测试`。日常以模块快检或模块局部数据库测试为完成证据,无证据严禁声称“完成/修复/通过”。
3. **数据权威与先审后入**:数据库为唯一正式权威,严禁裸连操作;正文、规划与知识抽取默认生成 Shadow 候选,经用户明确确认后方可写入 Canonical 正典事实。 3. **数据权威与先审后入**:数据库为唯一正式权威,严禁裸连操作;正文、规划与知识抽取默认生成 Shadow 候选,经用户明确确认后方可写入 Canonical 正典事实。
4. **模型治理与受控探索**:模型调用遵守[预算管理](src/muse/任务运行/预算管理.py)的固定日界窗口(默认 Asia/Shanghai,每日00/05/10/15/20开始,末窗20至24为4小时)与受控治理链;角色允许模型与策略版本以[角色策略](配置/角色策略.yaml)为准,并遵守[角色合同](.agent/角色/角色合同.md)(写手/规划固定顶级推理模型,裁判使用独立精确白名单);确定性逻辑、门禁与报告组装由脚本完成,严禁调用模型;智能体探索仅限圈定只读工具并留痕。 4. **模型治理与受控探索**:模型调用遵守[预算管理](src/muse/任务运行/预算管理.py)的固定日界窗口(默认 Asia/Shanghai,每日00/05/10/15/20开始,末窗20至24为4小时)与受控治理链;角色允许模型与策略版本以[角色策略](配置/角色策略.yaml)为准,并遵守[角色合同](.agent/角色/角色合同.md)(写手/规划固定顶级推理模型,裁判使用独立精确白名单);确定性逻辑、门禁与报告组装由脚本完成,严禁调用模型;智能体探索仅限圈定只读工具并留痕。
5. **会话交互与汇报纪律**:全程使用简体中文白话,坚决去除 AI 味(直陈事实、动作与后果,禁止清嗓子套话与空转缓冲词);需要用户决策时,必须交代清楚前因后果及各选项对下游的影响。 5. **会话交互与汇报纪律**:全程使用简体中文白话,坚决去除 AI 味(直陈事实、动作与后果,禁止清嗓子套话与空转缓冲词);需要用户决策时,必须交代清楚前因后果及各选项对下游的影响。

View File

@ -7,7 +7,7 @@ export PYTHONPATH := $(CURDIR)/src:$(CURDIR)/工具
# Muse 安装、生成、检查、测试和构建入口。 # Muse 安装、生成、检查、测试和构建入口。
# 每个声明命令都必须真实可执行(见 .agent/rules/文档与资源生成.md)。 # 每个声明命令都必须真实可执行(见 .agent/rules/文档与资源生成.md)。
.PHONY: 迁移 前端旅程 快检 索引生成 包验收 安装 格式 格式写入 类型 模块边界 检查 测试 数据库测试 数据库分片 验收数据库 评测实验 浏览器测试 宿主测试 真实模型测试 前端安装 前端检查 前端测试 生成 构建 资源核对 .PHONY: 迁移 前端旅程 快检 索引生成 包验收 安装 格式 格式写入 类型 模块边界 检查 测试 数据库测试 数据库分片 验收数据库 浏览器测试 宿主测试 真实模型测试 前端安装 前端检查 前端测试 生成 构建 资源核对
# 安装:建立并锁定 Python 依赖环境 # 安装:建立并锁定 Python 依赖环境
安装: 安装:
@ -65,7 +65,7 @@ export PYTHONPATH := $(CURDIR)/src:$(CURDIR)/工具
$(仓内解释器) -m pytest -m "not 数据库 and not 网络 and not 真实模型 and not 浏览器 and not 宿主 and not 安装包" $(范围) $(if $(用例),--case $(用例),) $(仓内解释器) -m pytest -m "not 数据库 and not 网络 and not 真实模型 and not 浏览器 and not 宿主 and not 安装包" $(范围) $(if $(用例),--case $(用例),)
# 验收数据库:全量高成本门禁,严禁自动执行,必须由用户显式提供 MUSE_RUN_FULL_DB=1。 # 验收数据库:全量高成本门禁,严禁自动执行,必须由用户显式提供 MUSE_RUN_FULL_DB=1。
# 先生成、再门禁、最后跑库;排除慢评测实验(慢评测由 make 评测实验 显式执行)。 # 先生成、再门禁、最后跑库;排除标记为慢的用例(浏览器旅程等走对应显式目标)。
验收数据库: 验收数据库:
@test -n "$$MUSE_TEST_DATABASE_URL" || { echo "缺少 MUSE_TEST_DATABASE_URL(隔离库连接串);拒绝先付生成与门禁成本再失败" >&2; exit 1; } @test -n "$$MUSE_TEST_DATABASE_URL" || { echo "缺少 MUSE_TEST_DATABASE_URL(隔离库连接串);拒绝先付生成与门禁成本再失败" >&2; exit 1; }
@test "$$MUSE_RUN_FULL_DB" = "1" || { echo "全量数据库验收耗时高且写盘大,必须用户显式授权。请传入 MUSE_RUN_FULL_DB=1 确认执行" >&2; exit 1; } @test "$$MUSE_RUN_FULL_DB" = "1" || { echo "全量数据库验收耗时高且写盘大,必须用户显式授权。请传入 MUSE_RUN_FULL_DB=1 确认执行" >&2; exit 1; }
@ -79,7 +79,7 @@ export PYTHONPATH := $(CURDIR)/src:$(CURDIR)/工具
# 数据库测试:需要 MUSE_TEST_DATABASE_URL。 # 数据库测试:需要 MUSE_TEST_DATABASE_URL。
# 未指定范围时的全量测试必须显式传入 MUSE_RUN_FULL_DB=1;指定范围(如 范围=tests/集成/test_某.py)可直接运行。 # 未指定范围时的全量测试必须显式传入 MUSE_RUN_FULL_DB=1;指定范围(如 范围=tests/集成/test_某.py)可直接运行。
# 默认排除慢评测实验、浏览器、网络、宿主、真实模型。 # 默认排除慢用例、浏览器、网络、宿主、真实模型。
数据库测试: 数据库测试:
@test -n "$$MUSE_TEST_DATABASE_URL" || { echo "缺少 MUSE_TEST_DATABASE_URL(隔离库连接串);拒绝静默跳过" >&2; exit 1; } @test -n "$$MUSE_TEST_DATABASE_URL" || { echo "缺少 MUSE_TEST_DATABASE_URL(隔离库连接串);拒绝静默跳过" >&2; exit 1; }
@if [ -z "$(范围)" ] && [ "$$MUSE_RUN_FULL_DB" != "1" ]; then \ @if [ -z "$(范围)" ] && [ "$$MUSE_RUN_FULL_DB" != "1" ]; then \
@ -87,11 +87,6 @@ export PYTHONPATH := $(CURDIR)/src:$(CURDIR)/工具
fi fi
$(仓内解释器) -m pytest --外部环境 -m "数据库 and not 慢 and not 宿主 and not 真实模型 and not 浏览器 and not 网络" $(范围) $(if $(用例),--case $(用例),) $(仓内解释器) -m pytest --外部环境 -m "数据库 and not 慢 and not 宿主 and not 真实模型 and not 浏览器 and not 网络" $(范围) $(if $(用例),--case $(用例),)
# 评测实验:显式运行大 N 慢用例(85次模型调用的资格终态);默认从所有日常套件中排除。
评测实验:
@test -n "$$MUSE_TEST_DATABASE_URL" || { echo "缺少 MUSE_TEST_DATABASE_URL(隔离库连接串)" >&2; exit 1; }
$(仓内解释器) -m pytest --外部环境 -m "数据库 and 慢" --timeout=900 $(范围)
# 浏览器测试:需要隔离库与显式浏览器路径;11 个 数据库+浏览器 用例的唯一入口。 # 浏览器测试:需要隔离库与显式浏览器路径;11 个 数据库+浏览器 用例的唯一入口。
# MUSE_ISOLATED_TEST_ENVIRONMENT 由本目标声明:Playwright 配置加载期就要求它,缺了连 spec 都跑不到。 # MUSE_ISOLATED_TEST_ENVIRONMENT 由本目标声明:Playwright 配置加载期就要求它,缺了连 spec 都跑不到。
浏览器测试: 浏览器测试:

View File

@ -23405,8 +23405,7 @@
"then": [ "then": [
"真实浏览器从同一报告读回,刷新不追加模型调用", "真实浏览器从同一报告读回,刷新不追加模型调用",
"资格报告显示实际标定覆盖和来源链接,读取不新增模型调用或启用", "资格报告显示实际标定覆盖和来源链接,读取不新增模型调用或启用",
"实际检测发现和原字引文在报告页读回,未知与未配置分别显示", "实际检测发现和原字引文在报告页读回,未知与未配置分别显示"
"实际效果分层、判据理由与凭据经工作台读回和封存;正式启用仍未评定"
], ],
"contract": "docs/系统架构/新版设计/模块设计/B10-效果评测.md", "contract": "docs/系统架构/新版设计/模块设计/B10-效果评测.md",
"file": "tests/集成/test_评测完整报告.py", "file": "tests/集成/test_评测完整报告.py",
@ -23416,7 +23415,6 @@
"calibration-执行环境5", "calibration-执行环境5",
"corrected-执行环境3", "corrected-执行环境3",
"detection-执行环境7", "detection-执行环境7",
"effect-执行环境8",
"judgment-执行环境2", "judgment-执行环境2",
"literary-执行环境4", "literary-执行环境4",
"plain-执行环境0", "plain-执行环境0",
@ -23427,7 +23425,6 @@
"tests/集成/test_评测完整报告.py::test_浏览器报告来自同一隔离实验且读取不追加调用__254004[calibration-执行环境5]", "tests/集成/test_评测完整报告.py::test_浏览器报告来自同一隔离实验且读取不追加调用__254004[calibration-执行环境5]",
"tests/集成/test_评测完整报告.py::test_浏览器报告来自同一隔离实验且读取不追加调用__254004[corrected-执行环境3]", "tests/集成/test_评测完整报告.py::test_浏览器报告来自同一隔离实验且读取不追加调用__254004[corrected-执行环境3]",
"tests/集成/test_评测完整报告.py::test_浏览器报告来自同一隔离实验且读取不追加调用__254004[detection-执行环境7]", "tests/集成/test_评测完整报告.py::test_浏览器报告来自同一隔离实验且读取不追加调用__254004[detection-执行环境7]",
"tests/集成/test_评测完整报告.py::test_浏览器报告来自同一隔离实验且读取不追加调用__254004[effect-执行环境8]",
"tests/集成/test_评测完整报告.py::test_浏览器报告来自同一隔离实验且读取不追加调用__254004[judgment-执行环境2]", "tests/集成/test_评测完整报告.py::test_浏览器报告来自同一隔离实验且读取不追加调用__254004[judgment-执行环境2]",
"tests/集成/test_评测完整报告.py::test_浏览器报告来自同一隔离实验且读取不追加调用__254004[literary-执行环境4]", "tests/集成/test_评测完整报告.py::test_浏览器报告来自同一隔离实验且读取不追加调用__254004[literary-执行环境4]",
"tests/集成/test_评测完整报告.py::test_浏览器报告来自同一隔离实验且读取不追加调用__254004[plain-执行环境0]", "tests/集成/test_评测完整报告.py::test_浏览器报告来自同一隔离实验且读取不追加调用__254004[plain-执行环境0]",
@ -27501,48 +27498,9 @@
"数据库" "数据库"
] ]
}, },
{
"case_id": "NC-w25-25f901",
"environment": "隔离PG;85次合成HTTP及测试runtime验证器,不认证真实外部模型效果",
"given": "确切方法确认版本和明确作者身份",
"when": "核对启用凭据并执行当前目标的作者状态操作",
"then": [
"原S02标定与留出通过后最小凭据导出、作者确认、按用途消费、原标定停止后拒绝新消费且保留历史"
],
"contract": "docs/系统架构/新版设计/模块设计/B10-效果评测.md",
"file": "tests/集成/test_方法启用凭据.py",
"symbol": "test_实际方法凭据作者启用及标定停止阻断新消费__25f901",
"parameter_ids": [
"执行环境0"
],
"node_ids": [
"tests/集成/test_方法启用凭据.py::test_实际方法凭据作者启用及标定停止阻断新消费__25f901[执行环境0]"
],
"fixtures": [
"monkeypatch",
"request",
"tmp_path",
"tmp_path_factory",
"内置种子方案",
"内置结构测试库",
"执行环境",
"数据库底座",
"方法环境",
"测试资源接缝",
"源码资源",
"离线防护",
"隔离数据库URL"
],
"markers": [
"parametrize",
"timeout",
"慢",
"数据库"
]
},
{ {
"case_id": "NC-w25-25f902", "case_id": "NC-w25-25f902",
"environment": "隔离PG;85次合成HTTP及测试runtime验证器,不认证真实外部模型效果", "environment": "隔离PG;无效凭据与角色权限,不发起合成评测循环",
"given": "确切方法确认版本和明确作者身份", "given": "确切方法确认版本和明确作者身份",
"when": "核对启用凭据并执行当前目标的作者状态操作", "when": "核对启用凭据并执行当前目标的作者状态操作",
"then": [ "then": [
@ -28259,45 +28217,6 @@
"数据库" "数据库"
] ]
}, },
{
"case_id": "NC-w26-26c001",
"environment": "隔离PG、实际B10/S01/B04与合成HTTP;不认证外部文学效果",
"given": "确切方法版本、真实反馈和同作者用途配置",
"when": "登记具体改进、实际验证和owner启停并恢复已提交结果;记录分版本观察",
"then": [
"具体改进实际评测与owner回执中断恢复"
],
"contract": "docs/系统架构/新版设计/模块设计/B07-作者经验.md",
"file": "tests/集成/test_具体改进与回执.py",
"symbol": "test_具体改进实际评测与owner回执中断恢复__26c001",
"parameter_ids": [
"执行环境0"
],
"node_ids": [
"tests/集成/test_具体改进与回执.py::test_具体改进实际评测与owner回执中断恢复__26c001[执行环境0]"
],
"fixtures": [
"monkeypatch",
"request",
"tmp_path",
"tmp_path_factory",
"内置种子方案",
"内置结构测试库",
"执行环境",
"数据库底座",
"方法环境",
"测试资源接缝",
"源码资源",
"离线防护",
"隔离数据库URL"
],
"markers": [
"parametrize",
"timeout",
"慢",
"数据库"
]
},
{ {
"case_id": "NC-w26-26c002", "case_id": "NC-w26-26c002",
"environment": "隔离PG、实际B10/S01/B04与合成HTTP;不认证外部文学效果", "environment": "隔离PG、实际B10/S01/B04与合成HTTP;不认证外部文学效果",
@ -33163,46 +33082,6 @@
"数据库" "数据库"
] ]
}, },
{
"case_id": "TC-0fe403cb0320",
"environment": "隔离PG、实际S02与合成HTTP;不认证外部模型文学收益",
"given": "本例固定样本、场景与独立异常,不共享其他用例的执行结果",
"when": "经真实生成、独立检测、比较与公开效果入口读回",
"then": [
"真实合格标定后的无增益留出集返回no_gain",
"封存拒绝且公开凭据为空",
"隔离PG效果凭据表无记录,代替旧文件不存在断言"
],
"contract": "docs/系统架构/新版设计/模块设计/B10-效果评测.md",
"file": "tests/集成/test_实际效果判据.py",
"symbol": "test_真实无增益留出实验不产生合格凭据__25f406",
"parameter_ids": [
"执行环境0"
],
"node_ids": [
"tests/集成/test_实际效果判据.py::test_真实无增益留出实验不产生合格凭据__25f406[执行环境0]"
],
"fixtures": [
"monkeypatch",
"request",
"tmp_path",
"tmp_path_factory",
"内置种子方案",
"内置结构测试库",
"执行环境",
"数据库底座",
"测试资源接缝",
"源码资源",
"离线防护",
"隔离数据库URL"
],
"markers": [
"parametrize",
"timeout",
"慢",
"数据库"
]
},
{ {
"case_id": "TC-1062c87c6b5e", "case_id": "TC-1062c87c6b5e",
"environment": "隔离 PostgreSQL,检测替身", "environment": "隔离 PostgreSQL,检测替身",
@ -33344,47 +33223,6 @@
], ],
"markers": [] "markers": []
}, },
{
"case_id": "TC-1569749609e7",
"environment": "隔离PG、实际S02与合成HTTP;不认证外部模型文学收益",
"given": "本例固定样本、场景与独立异常,不共享其他用例的执行结果",
"when": "经真实生成、独立检测、比较与公开效果入口读回",
"then": [
"实际标定与两作品留出集效果passed",
"不可变PG凭据存在并关联原目标",
"重复读回及HTTP、CLI与原凭据哈希一致",
"对派生凭据篡改且重签也因逐例重算不符被拒绝"
],
"contract": "docs/系统架构/新版设计/模块设计/B10-效果评测.md",
"file": "tests/集成/test_实际效果判据.py",
"symbol": "test_完整原标定和留出实验通过并封存可复算凭据__25f404",
"parameter_ids": [
"执行环境0"
],
"node_ids": [
"tests/集成/test_实际效果判据.py::test_完整原标定和留出实验通过并封存可复算凭据__25f404[执行环境0]"
],
"fixtures": [
"monkeypatch",
"request",
"tmp_path",
"tmp_path_factory",
"内置种子方案",
"内置结构测试库",
"执行环境",
"数据库底座",
"测试资源接缝",
"源码资源",
"离线防护",
"隔离数据库URL"
],
"markers": [
"parametrize",
"timeout",
"慢",
"数据库"
]
},
{ {
"case_id": "TC-15c187553e39", "case_id": "TC-15c187553e39",
"environment": "隔离PostgreSQL与真实S02/业务owner,模型为明确合成HTTP", "environment": "隔离PostgreSQL与真实S02/业务owner,模型为明确合成HTTP",

View File

@ -5,14 +5,12 @@ import os
import subprocess import subprocess
import sys import sys
from dataclasses import asdict, replace from dataclasses import asdict, replace
from datetime import UTC, datetime, timedelta
from uuid import uuid4 from uuid import uuid4
import pytest import pytest
import test_方法启用凭据 as 启用测试 import test_方法启用凭据 as 启用测试
import test_正文候选与作者决定 as 正文测试 import test_正文候选与作者决定 as 正文测试
from muse.作者经验.存储 import 经验存储
from muse.作者经验.接口 import 反馈版本, 反馈请求, 改进定位, 改进请求 from muse.作者经验.接口 import 反馈版本, 反馈请求, 改进定位, 改进请求
from muse.共享.调用身份 import 用途 from muse.共享.调用身份 import 用途
from muse.共享.错误 import Muse错误 from muse.共享.错误 import Muse错误
@ -57,99 +55,6 @@ def _提案(env):
return service, request, 改进定位(request.improvement_id, row["revision"], row["content_hash"]) return service, request, 改进定位(request.improvement_id, row["revision"], row["content_hash"])
@pytest.mark.case_id(
"NC-w26-26c001",
environment="隔离PG、实际B10/S01/B04与合成HTTP;不认证外部文学效果",
given="确切方法版本、真实反馈和同作者用途配置",
when="登记具体改进、实际验证和owner启停并恢复已提交结果;记录分版本观察",
then=["具体改进实际评测与owner回执中断恢复"],
contract="docs/系统架构/新版设计/模块设计/B07-作者经验.md",
)
@pytest.mark.parametrize("执行环境", [启用测试.参数], indirect=True)
@pytest.mark.timeout(900)
@pytest.mark.慢
def test_具体改进实际评测与owner回执中断恢复__26c001(执行环境, 方法环境, monkeypatch):
env = 启用测试._确认方法(执行环境, 方法环境)
service, request, point = _提案(env)
author = env["owner"]
evaluation = env["app"].要求评测()
certificate = 启用测试.效果测试._校准(env)
write_result = 经验存储.保存改进结果
failures = {"validate": 1, "enable": 1}
def 断一次(self, who, command, receipt):
if failures.get(command, 0):
failures[command] -= 1
raise RuntimeError("目标已经提交,登记前中断")
return write_result(self, who, command, receipt)
monkeypatch.setattr(经验存储, "保存改进结果", 断一次)
def 注册(req):
with pytest.raises(RuntimeError, match="登记前中断"):
service.申请验证(author, "validate", point, evaluation, env["actor"], req)
pending = service.读取改进(author, request.improvement_id)["actions"][0]
assert pending["receipt"] is None
actual = evaluation.查找已登记实验(
env["actor"], "B07.validation:" + str(pending["action_id"]), req
)
assert actual is not None
recovered = service.申请验证(author, "validate", point, evaluation, env["actor"], req)
assert recovered["receipt"]["experiment_id"] == actual["experiment_id"]
assert set(recovered["receipt"]) == {
"experiment_id",
"conditions_hash",
"target",
"request_hash",
"registration",
}
return actual
dest = 启用测试._实际方法留出(env, certificate, 登记=注册)
assert 启用测试.效果测试._完成(dest)["assessment"]["gate_b"]["status"] == "passed"
eid = dest["exp"]["experiment_id"]
evaluation.封存效果判据(env["actor"], eid)
proof = evaluation.导出启用凭据(
env["actor"],
eid,
env["pools"][用途.维护],
replace(env["actor"], 用途=用途.维护),
批准引用="synthetic-improvement-only",
有效期=datetime.now(UTC) + timedelta(hours=1),
)
with pytest.raises(RuntimeError, match="登记前中断"):
service.请求启停(author, "enable", point, "enable", proof["receipt_id"])
after = service.读取改进(author, request.improvement_id)
assert after["target_state"] == "enabled"
intent = next(x for x in after["actions"] if x["kind"] == "enable")
assert intent["receipt"] is None
original = service.正式.读取回执(author, "B07.owner:" + str(intent["action_id"]))
evaluation.取消实验(
env["actor"], certificate["result"]["experiment_id"], "stop-before-recovery"
)
recovered = service.请求启停(author, "enable", point, "enable", proof["receipt_id"])
assert recovered["receipt"] == original
assert (
service.请求启停(author, "enable", point, "enable", proof["receipt_id"])["receipt"]
== original
)
disabled = service.请求启停(author, "disable", point, "disable")
assert disabled["receipt"]["results"][0]["state"] == "disabled"
assert service.读取改进(author, request.improvement_id)["target_state"] == "disabled"
with pytest.raises(Muse错误):
service.请求启停(author, "reenable", point, "enable", proof["receipt_id"])
service.提出改进(
author, replace(request, expected_revision=1, rationale="改进理由已由作者修正。")
)
with pytest.raises(Muse错误, match="当前版本"):
service.请求启停(author, "old-approval", point, "enable", proof["receipt_id"])
assert (
service.请求启停(author, "enable", point, "enable", proof["receipt_id"])["receipt"]
== original
)
assert len(env["received"]) == 85
@pytest.mark.case_id( @pytest.mark.case_id(
"NC-w26-26c002", "NC-w26-26c002",
environment="隔离PG、实际B10/S01/B04与合成HTTP;不认证外部文学效果", environment="隔离PG、实际B10/S01/B04与合成HTTP;不认证外部文学效果",

View File

@ -10,7 +10,7 @@ import test_评测执行与失败收敛 as 运行测试
import test_评测语义检测 as 检测测试 import test_评测语义检测 as 检测测试
from muse.共享.调用身份 import 用途 from muse.共享.调用身份 import 用途
from muse.效果评测.接口 import 实验请求, 数据集发布, 评测服务, 评测错误, 金标准发布 from muse.效果评测.接口 import 实验请求, 数据集发布, 评测服务, 评测错误
pytestmark = pytest.mark.数据库 pytestmark = pytest.mark.数据库
执行环境 = 运行测试.执行环境 执行环境 = 运行测试.执行环境
@ -186,46 +186,6 @@ def _完成(env):
return env["app"].要求评测().读取效果判据(env["actor"], env["exp"]["experiment_id"]) return env["app"].要求评测().读取效果判据(env["actor"], env["exp"]["experiment_id"])
def _校准(env):
cal = _新实验(env, calibration=True)
control = env["fixture_control"]
control.update(calibrating=True, judge_index=0)
运行测试._启动执行(cal)
运行测试._运行就绪(cal)
annotations = []
for unit in 文学测试._读(cal)["units"]:
if unit["kind"] != "generation":
continue
score = _基分("\n".join(p["text"] for p in unit["output"]["paragraphs"]), control)
annotations.append(
{
"unit_id": unit["unit_id"],
"output_hash": unit["evidence"]["structured_output_hash"],
"scores": {d: score for d in 文学测试.维度},
"assertions": {"guard": "pass"},
"constraints": {"stay": "pass"},
}
)
评测服务(env["pools"][用途.维护]).发布标定金标准(
replace(env["actor"], 用途=用途.维护),
金标准发布.model_validate(
{
"experiment_id": cal["exp"]["experiment_id"],
"approval_ref": "synthetic:fixed-labels",
"annotations": annotations,
}
),
)
文学测试._推进(cal)
运行测试._运行就绪(cal)
文学测试._推进(cal)
运行测试._运行就绪(cal)
cert = env["app"].要求评测().封存标定(env["actor"], cal["exp"]["experiment_id"])
assert cert["result"]["status"] == "passed" and control["judge_index"] == 15
control["calibrating"] = False
return cert
@pytest.mark.case_id( @pytest.mark.case_id(
"TC-64bcf0ac927f", "TC-64bcf0ac927f",
environment="隔离PG、实际S02与合成HTTP;不认证外部模型文学收益", environment="隔离PG、实际S02与合成HTTP;不认证外部模型文学收益",
@ -298,120 +258,6 @@ def test_单作品完整A阶段仍保留选择偏差__25f403(执行环境):
assert result["assessment"]["gate_b"]["status"] == "insufficient_evidence" assert result["assessment"]["gate_b"]["status"] == "insufficient_evidence"
@pytest.mark.case_id(
"TC-1569749609e7",
environment="隔离PG、实际S02与合成HTTP;不认证外部模型文学收益",
given="本例固定样本、场景与独立异常,不共享其他用例的执行结果",
when="经真实生成、独立检测、比较与公开效果入口读回",
then=[
"实际标定与两作品留出集效果passed",
"不可变PG凭据存在并关联原目标",
"重复读回及HTTP、CLI与原凭据哈希一致",
"对派生凭据篡改且重签也因逐例重算不符被拒绝",
],
contract="docs/系统架构/新版设计/模块设计/B10-效果评测.md",
)
@pytest.mark.timeout(900)
@pytest.mark.慢
@pytest.mark.parametrize("执行环境", [参数], indirect=True)
def test_完整原标定和留出实验通过并封存可复算凭据__25f404(执行环境, monkeypatch, tmp_path):
import copy
import subprocess
import sys
import psycopg
import test_生产评测权限隔离 as 数据测试
from fastapi.testclient import TestClient
import muse.效果评测.启用凭据 as 启用凭据模块
import muse.效果评测.启用判据 as effects
import muse.效果评测.接口 as 评测接口模块
from muse.接入.http.应用 import 创建应用
from muse.效果评测.存储 import 评测存储
from muse.正式变更.接口 import 固定哈希
from muse.配置 import 读取配置
env = 执行环境
cert = _校准(env)
dest = _新实验(env, 10, certificate=cert)
result = _完成(dest)
assert result["assessment"]["gate_b"]["status"] == "passed"
assert result["assessment"]["metrics"]["average_deltas"]["setting_entity_fidelity"] == 0.5
svc = env["app"].要求评测()
eid = dest["exp"]["experiment_id"]
policy_reader = effects.读取效果标准
def 改动后标准(version):
return {**policy_reader(version), "maximum_unstable_ratio": 0.19}
# 各消费模块在导入时直接绑定了该函数,改动要同时落到实际使用它的模块。
for 模块 in (effects, 启用凭据模块, 评测接口模块):
monkeypatch.setattr(模块, "读取效果标准", 改动后标准)
assert not svc.读取效果判据(env["actor"], eid)["current_policy"]
with pytest.raises(评测错误, match="标准过期"):
svc.封存效果判据(env["actor"], eid)
for 模块 in (effects, 启用凭据模块, 评测接口模块):
monkeypatch.setattr(模块, "读取效果标准", policy_reader)
sealed = svc.封存效果判据(env["actor"], eid)
assert sealed["receipt_id"] and svc.封存效果判据(env["actor"], eid) == sealed
assert sealed["assessment"]["target"] == dest["exp"]["conditions"]["target"]
assert sealed["receipt_hash"] == 固定哈希(sealed["assessment"])
with (
env["pools"][用途.生产].连接(只读=True) as conn,
pytest.raises(psycopg.errors.InsufficientPrivilege),
):
conn.execute("SELECT * FROM evaluation.muse_effect_assessment")
with env["pools"][用途.维护].连接() as conn, pytest.raises(psycopg.Error):
conn.execute("UPDATE evaluation.muse_effect_assessment SET payload=payload")
report = svc.读取实验报告(env["actor"], eid)
assert report["effect"] == sealed and report["activation_status"] == "not_evaluated"
assert sealed["assessment"]["validation_modes"] == ["offline_contract"]
config = 数据测试._配置文件(env["pool"], tmp_path)
http = 创建应用(读取配置(config))
with TestClient(http, headers={"origin": "http://testserver"}) as client:
http.state.装配 = env["app"]
assert (
client.post(
"/api/v1/session", json={"password": "synthetic-evaluation-only"}
).status_code
== 200
)
observed = client.get(f"/api/v1/evaluation/experiments/{eid}/effect")
assert observed.status_code == 200 and observed.json() == sealed
command = subprocess.run(
[sys.executable, "-I", "-m", "muse", "评测", str(config), "效果判据", eid],
cwd=tmp_path,
capture_output=True,
text=True,
timeout=45,
)
assert command.returncode == 0, command.stderr
assert json.loads(command.stdout) == sealed
with pytest.raises(评测错误):
svc.读取效果判据(replace(env["actor"], 作者="foreign-author"), eid)
read = 评测存储.读取效果凭据
def changed(self, experiment_id):
row = copy.deepcopy(read(self, experiment_id))
if row:
row["payload"]["gate_b"]["status"] = "failed"
row["payload_hash"] = 固定哈希(row["payload"])
return row
monkeypatch.setattr(评测存储, "读取效果凭据", changed)
with pytest.raises(评测错误, match="逐例重算"):
svc.读取效果判据(env["actor"], eid)
monkeypatch.setattr(评测存储, "读取效果凭据", read)
svc.取消实验(env["actor"], eid, "stop-after-effect")
old = svc.封存效果判据(env["actor"], eid)
assert old["stopped"] and old["assessment"] == sealed["assessment"]
svc.取消实验(env["actor"], cert["result"]["experiment_id"], "stop-source-after-effect")
historical = svc.读取效果判据(env["actor"], eid)
assert historical["calibration_stopped"] == [cert["result"]["experiment_id"]]
assert historical["assessment"] == sealed["assessment"]
assert len(env["received"]) == 85
@pytest.mark.case_id( @pytest.mark.case_id(
"NC-w25-25f405", "NC-w25-25f405",
environment="隔离PG与合成HTTP", environment="隔离PG与合成HTTP",
@ -440,42 +286,6 @@ def test_单次读取复用已核验交付而新读取仍重验__25f405(执行
assert len(calls) == 30 and set(calls.values()) == {1} assert len(calls) == 30 and set(calls.values()) == {1}
@pytest.mark.case_id(
"TC-0fe403cb0320",
environment="隔离PG、实际S02与合成HTTP;不认证外部模型文学收益",
given="本例固定样本、场景与独立异常,不共享其他用例的执行结果",
when="经真实生成、独立检测、比较与公开效果入口读回",
then=[
"真实合格标定后的无增益留出集返回no_gain",
"封存拒绝且公开凭据为空",
"隔离PG效果凭据表无记录,代替旧文件不存在断言",
],
contract="docs/系统架构/新版设计/模块设计/B10-效果评测.md",
)
@pytest.mark.timeout(900)
@pytest.mark.慢
@pytest.mark.parametrize("执行环境", [参数], indirect=True)
def test_真实无增益留出实验不产生合格凭据__25f406(执行环境):
env = 执行环境
env["fixture_control"]["gain"] = False
cert = _校准(env)
dest = _新实验(env, 10, certificate=cert)
result = _完成(dest)
assert result["assessment"]["gate_b"]["status"] == "no_gain"
with pytest.raises(评测错误, match="效果未通过"):
env["app"].要求评测().封存效果判据(env["actor"], dest["exp"]["experiment_id"])
assert result["receipt_id"] is None and len(env["received"]) == 85
assert (
env["app"].要求评测().读取效果判据(env["actor"], dest["exp"]["experiment_id"])["receipt_id"]
is None
)
with env["pool"].连接(只读=True) as conn:
assert (
conn.execute("SELECT count(*) FROM evaluation.muse_effect_assessment").fetchone()[0]
== 0
)
@pytest.mark.case_id( @pytest.mark.case_id(
"TC-602d0bb37aa2", "TC-602d0bb37aa2",
environment="隔离PG、实际S02与合成HTTP;不认证外部模型文学收益", environment="隔离PG、实际S02与合成HTTP;不认证外部模型文学收益",

View File

@ -4,22 +4,18 @@ runtime配置验证由测试夹具模拟,用来走有证据的分支;它不
""" """
import json import json
import os
from dataclasses import replace from dataclasses import replace
from datetime import UTC, datetime, timedelta
from uuid import uuid4 from uuid import uuid4
import psycopg import psycopg
import pytest import pytest
import test_上下文快照与索引写入 as 方法测试 import test_上下文快照与索引写入 as 方法测试
import test_实际效果判据 as 效果测试 import test_实际效果判据 as 效果测试
import test_生产评测权限隔离 as 数据测试
import test_评测执行与失败收敛 as 运行测试 import test_评测执行与失败收敛 as 运行测试
from muse.共享.调用身份 import 用途 from muse.共享.调用身份 import 用途
from muse.共享.错误 import Muse错误 from muse.共享.错误 import Muse错误
from muse.效果评测.接口 import 实验请求, 方法数据集发布, 评测服务, 评测错误, 读取启用凭据 from muse.知识方法.接口 import 状态命令
from muse.知识方法.接口 import 状态命令, 绑定命令, 读取绑定材料
pytestmark = pytest.mark.数据库 pytestmark = pytest.mark.数据库
执行环境 = 运行测试.执行环境 执行环境 = 运行测试.执行环境
@ -65,199 +61,9 @@ def _确认方法(env, method_env):
} }
def _实际方法留出(env, certificate, 登记=None):
svc = 评测服务(env["pools"][用途.维护])
data = svc.发布方法数据集(
replace(env["actor"], 用途=用途.维护),
方法数据集发布.model_validate(
{
"dataset_id": "enablement-method-holdout",
"revision": 1,
"samples": 效果测试._数据样本(10),
"method_version_id": str(env["version"]["version_id"]),
"approval_ref": "synthetic:method-material",
"expires_at": datetime.now(UTC) + timedelta(hours=1),
}
),
)
req = 实验请求.model_validate(
{
**env["request"].model_dump(mode="json"),
"dataset_version_id": data["version_id"],
"dataset_hash": data["public_hash"],
"target": data["method_target"],
"split": "holdout",
"max_cost_usd": "20",
"evaluation_goal": "qualification",
"effect_policy": "writer-effect-v1",
"calibration_use": {
"policy": 效果测试.标定策略,
"references": [
{
"experiment_id": certificate["result"]["experiment_id"],
"receipt_hash": certificate["receipt_hash"],
}
],
},
}
)
exp = (
登记(req)
if 登记
else env["app"].要求评测().创建实验(env["actor"], "enablement-holdout", req)
)
return {**env, "exp": exp, "request": req}
@pytest.mark.case_id(
"NC-w25-25f901",
environment="隔离PG;85次合成HTTP及测试runtime验证器,不认证真实外部模型效果",
given="确切方法确认版本和明确作者身份",
when="核对启用凭据并执行当前目标的作者状态操作",
then=[
"原S02标定与留出通过后最小凭据导出、作者确认、按用途消费、原标定停止后拒绝新消费且保留历史"
],
contract="docs/系统架构/新版设计/模块设计/B10-效果评测.md",
)
@pytest.mark.parametrize("执行环境", [参数], indirect=True)
@pytest.mark.timeout(900)
@pytest.mark.慢
def test_实际方法凭据作者启用及标定停止阻断新消费__25f901(执行环境, 方法环境, tmp_path):
from fastapi.testclient import TestClient
from muse.接入.http.应用 import 创建应用
from muse.配置 import 读取配置
env = _确认方法(执行环境, 方法环境)
certificate = 效果测试._校准(env)
dest = _实际方法留出(env, certificate)
assert 效果测试._完成(dest)["assessment"]["gate_b"]["status"] == "passed"
svc, eid = env["app"].要求评测(), dest["exp"]["experiment_id"]
sealed = svc.封存效果判据(env["actor"], eid)
assert sealed["assessment"]["validation_modes"] == ["runtime"]
now = datetime.now(UTC)
arguments = dict(批准引用="synthetic:export-proof", 有效期=now + timedelta(hours=2))
def export():
return svc.导出启用凭据(
env["actor"],
eid,
env["pools"][用途.维护],
replace(env["actor"], 用途=用途.维护),
**arguments,
)
proof = export()
assert export() == proof
rid = proof["receipt_id"]
assert proof["payload"]["target"] == dest["exp"]["conditions"]["target"]
assert not any(
word in json.dumps(proof, ensure_ascii=False)
for word in [
"paragraphs",
"candidate_scores",
"ORACLE-",
"left",
"right",
"source_authorization",
]
)
methods, owner, mid = env["methods"], env["owner"], env["method_id"]
assert methods.读取方法详情(owner, mid)["state"] == "confirmed"
assert methods.核对启用(owner, mid, rid)["receipt_hash"] == proof["receipt_hash"]
for purpose in (用途.生产, 用途.评测):
with (
env["pools"][purpose].连接() as conn,
pytest.raises(psycopg.errors.InsufficientPrivilege),
):
conn.execute(
"INSERT INTO public.muse_enablement_receipt "
"SELECT * FROM public.muse_enablement_receipt"
)
with (
env["pools"][用途.生产].连接() as conn,
pytest.raises(psycopg.errors.InsufficientPrivilege),
):
conn.execute("SELECT * FROM oracle.muse_dataset_answers")
with env["pools"][用途.生产].连接(只读=True) as conn:
assert 读取启用凭据(conn, owner.作者, rid) == proof
with pytest.raises(Muse错误):
methods.核对启用(replace(owner, 作者="wrong-owner"), mid, rid)
config = 数据测试._配置文件(env["pools"][用途.生产], tmp_path)
config.write_text(config.read_text().replace("eval-author", owner.作者))
browser_receipt = None
if os.environ.get("MUSE_RUN_BROWSER") == "1":
browser_receipt = _浏览器启用(env, config, rid, tmp_path)
with TestClient(创建应用(读取配置(config)), headers={"origin": "http://testserver"}) as client:
assert (
client.post(
"/api/v1/session", json={"password": "synthetic-evaluation-only"}
).status_code
== 200
)
assert client.get(f"/api/v1/methods/{mid}/enablement/{rid}").status_code == 200
body = {
"command_id": browser_receipt["command_id"]
if browser_receipt
else "enable-tested-method",
"method_id": mid,
"action": "enable",
"启用凭据": rid,
"expected_state_revision": 1,
}
assert client.post(f"/api/v1/methods/{uuid4()}/state", json=body).status_code == 403
adopted = client.post(f"/api/v1/methods/{mid}/state", json=body)
assert adopted.status_code == 200, adopted.text
assert adopted.json()["results"][0]["state"] == "enabled"
if browser_receipt:
assert adopted.json() == browser_receipt
assert client.post(f"/api/v1/methods/{mid}/state", json=body).json() == adopted.json()
methods.绑定方法(
owner,
"bind-enabled-method",
绑定命令(
"method-work",
"work",
"work",
mid,
str(env["version"]["version_id"]),
),
)
with env["pools"][用途.生产].连接(只读=True) as conn:
material = 读取绑定材料(conn, "method-work", 内容用途="generation", 运行用途="production")
assert len(material) == 1
# 写手凭据不能扩为规划用途;拒绝码由原来的用途复核前移到凭据范围守卫。
with pytest.raises(Muse错误, match="不能扩为新目标或其他用途"):
读取绑定材料(conn, "method-work", 内容用途="planning", 运行用途="production")
usages = methods.列出消费(owner, str(env["version"]["version_id"]))
assert len(usages) == 1 and usages[0]["kind"] == "enablement"
assert usages[0]["result_ref"] == "B10.enablement:" + rid
before = len(env["received"])
svc.取消实验(env["actor"], certificate["result"]["experiment_id"], "stop-calibration")
with env["pools"][用途.生产].连接(只读=True) as conn:
assert 读取启用凭据(conn, owner.作者, rid)["stopped"]
with pytest.raises(Muse错误, match="停止"):
读取绑定材料(conn, "method-work", 内容用途="generation", 运行用途="production")
methods.变更方法状态(
owner, "disable-tested-method", 状态命令(mid, "disable", expected_state_revision=1)
)
with pytest.raises(Muse错误, match="停止"):
methods.变更方法状态(
owner,
"reenable-stopped-method",
状态命令(mid, "enable", expected_state_revision=1, 启用凭据=rid),
)
with pytest.raises(评测错误):
export()
assert methods.列出消费(owner, str(env["version"]["version_id"])) == usages
assert len(env["received"]) == before == 85
@pytest.mark.case_id( @pytest.mark.case_id(
"NC-w25-25f902", "NC-w25-25f902",
environment="隔离PG;85次合成HTTP及测试runtime验证器,不认证真实外部模型效果", environment="隔离PG;无效凭据与角色权限,不发起合成评测循环",
given="确切方法确认版本和明确作者身份", given="确切方法确认版本和明确作者身份",
when="核对启用凭据并执行当前目标的作者状态操作", when="核对启用凭据并执行当前目标的作者状态操作",
then=["不存在或格式错误的凭据不能启用,生产和评测角色不能写凭据原表"], then=["不存在或格式错误的凭据不能启用,生产和评测角色不能写凭据原表"],
@ -294,97 +100,3 @@ def test_未有目标证明不能启用且公开表不可写__25f902(方法环
"SELECT has_table_privilege(current_user,%s,'SELECT')", "SELECT has_table_privilege(current_user,%s,'SELECT')",
("public.muse_enablement_status",), ("public.muse_enablement_status",),
).fetchone()[0] ).fetchone()[0]
def _浏览器启用(env, config, rid, tmp_path):
import socket
import subprocess
from contextlib import asynccontextmanager
from pathlib import Path
from threading import Event, Thread
import uvicorn
from muse.接入.http.应用 import 创建应用
from muse.配置 import 服务配置, 读取配置
password = tmp_path / "method-browser-key"
password.write_text("synthetic-evaluation-only")
password.chmod(0o600)
sock = socket.socket()
sock.bind(("127.0.0.1", 0))
port = sock.getsockname()[1]
base = f"http://127.0.0.1:{port}"
app = 创建应用(
replace(
读取配置(config),
HTTP=服务配置(
str(password),
作者ID=env["owner"].作者,
端口=port,
公开地址=base,
允许来源=(base,),
),
)
)
previous = app.router.lifespan_context
ready = Event()
@asynccontextmanager
async def lifespan(instance):
async with previous(instance):
ready.set()
yield
app.router.lifespan_context = lifespan
server = uvicorn.Server(
uvicorn.Config(
app,
log_level="warning",
access_log=False,
timeout_graceful_shutdown=3,
)
)
thread = Thread(target=server.run, kwargs={"sockets": [sock]}, daemon=True)
output = Path(os.environ.get("MUSE_BROWSER_ARTIFACTS", str(tmp_path / "browser"))).resolve()
output.mkdir(parents=True, exist_ok=True)
process_env = {
**os.environ,
"MUSE_WORKBENCH_URL": base,
"MUSE_AUTHOR_PASSWORD_FILE": str(password),
"MUSE_METHOD_ID": env["method_id"],
"MUSE_ENABLEMENT_RECEIPT": rid,
}
chrome = Path("/Applications/Google Chrome.app/Contents/MacOS/Google Chrome")
if chrome.exists():
process_env.setdefault("MUSE_BROWSER_EXECUTABLE", str(chrome))
thread.start()
try:
assert ready.wait(10)
with (output / "浏览器.json").open("w") as log:
result = subprocess.run(
[
"pnpm",
"exec",
"playwright",
"test",
"方法启用.spec.ts",
"--reporter=json",
"--output",
str(output / "产物"),
],
cwd=Path(__file__).parents[2] / "web",
env=process_env,
stdout=log,
stderr=subprocess.STDOUT,
timeout=90,
)
assert result.returncode == 0, str(output / "浏览器.json")
receipts = list(output.rglob("启用回执.json"))
assert len(receipts) == 1
return json.loads(receipts[0].read_text())
finally:
server.should_exit = True
thread.join(timeout=10)
sock.close()
assert not thread.is_alive()

View File

@ -4,7 +4,6 @@ from decimal import Decimal
import pytest import pytest
import test_ABC正文回放 as ABC测试 import test_ABC正文回放 as ABC测试
import test_实际效果判据 as 效果测试
import test_文学评分执行与条件第三 as 文学测试 import test_文学评分执行与条件第三 as 文学测试
import test_标定凭据消费 as 消费测试 import test_标定凭据消费 as 消费测试
import test_标定金标准与凭据 as 标定测试 import test_标定金标准与凭据 as 标定测试
@ -126,7 +125,6 @@ def test_独立评委分歧逐维保留且不生成启用许可__254003(执行
"真实浏览器从同一报告读回,刷新不追加模型调用", "真实浏览器从同一报告读回,刷新不追加模型调用",
"资格报告显示实际标定覆盖和来源链接,读取不新增模型调用或启用", "资格报告显示实际标定覆盖和来源链接,读取不新增模型调用或启用",
"实际检测发现和原字引文在报告页读回,未知与未配置分别显示", "实际检测发现和原字引文在报告页读回,未知与未配置分别显示",
"实际效果分层、判据理由与凭据经工作台读回和封存;正式启用仍未评定",
], ],
contract="docs/系统架构/新版设计/模块设计/B10-效果评测.md", contract="docs/系统架构/新版设计/模块设计/B10-效果评测.md",
) )
@ -144,7 +142,6 @@ def test_独立评委分歧逐维保留且不生成启用许可__254003(执行
("calibration", 标定测试.参数), ("calibration", 标定测试.参数),
("qualification", 标定测试.参数), ("qualification", 标定测试.参数),
("detection", 检测测试.参数), ("detection", 检测测试.参数),
("effect", 效果测试.参数),
], ],
indirect=["执行环境"], indirect=["执行环境"],
) )
@ -168,9 +165,6 @@ def test_浏览器报告来自同一隔离实验且读取不追加调用__254004
if os.environ.get("MUSE_RUN_BROWSER") != "1": if os.environ.get("MUSE_RUN_BROWSER") != "1":
pytest.skip("浏览器未显式启用,不计为通过") pytest.skip("浏览器未显式启用,不计为通过")
env = request.getfixturevalue("ABC环境") if report_kind == "abc" else 执行环境 env = request.getfixturevalue("ABC环境") if report_kind == "abc" else 执行环境
if report_kind == "effect":
cert = 效果测试._校准(env)
env = 效果测试._新实验(env, 10, certificate=cert)
if report_kind == "qualification": if report_kind == "qualification":
cert = 消费测试._校准(env) cert = 消费测试._校准(env)
env = 消费测试._新实验(env, 消费测试._请求(env, [cert])) env = 消费测试._新实验(env, 消费测试._请求(env, [cert]))
@ -181,7 +175,6 @@ def test_浏览器报告来自同一隔离实验且读取不追加调用__254004
"calibration", "calibration",
"qualification", "qualification",
"detection", "detection",
"effect",
}: }:
env["scripted"].extend(["ok", "unknown_cost"]) env["scripted"].extend(["ok", "unknown_cost"])
运行测试._启动执行(env) 运行测试._启动执行(env)
@ -195,7 +188,6 @@ def test_浏览器报告来自同一隔离实验且读取不追加调用__254004
"calibration", "calibration",
"qualification", "qualification",
"detection", "detection",
"effect",
}: }:
env["app"].要求评测().推进实验(env["actor"], env["exp"]["experiment_id"]) env["app"].要求评测().推进实验(env["actor"], env["exp"]["experiment_id"])
if report_kind == "corrected": if report_kind == "corrected":
@ -205,7 +197,7 @@ def test_浏览器报告来自同一隔离实验且读取不追加调用__254004
if report_kind == "detection": if report_kind == "detection":
env["scripted"].extend(["high", "ok"]) env["scripted"].extend(["high", "ok"])
运行测试._运行就绪(env) 运行测试._运行就绪(env)
if report_kind in {"literary", "calibration", "qualification", "detection", "effect"}: if report_kind in {"literary", "calibration", "qualification", "detection"}:
env["app"].要求评测().推进实验(env["actor"], env["exp"]["experiment_id"]) env["app"].要求评测().推进实验(env["actor"], env["exp"]["experiment_id"])
运行测试._运行就绪(env) 运行测试._运行就绪(env)
before = _报告(env) before = _报告(env)
@ -260,7 +252,6 @@ def test_浏览器报告来自同一隔离实验且读取不追加调用__254004
"calibration", "calibration",
"qualification", "qualification",
"detection", "detection",
"effect",
} }
else "0", else "0",
} }
@ -299,10 +290,6 @@ def test_浏览器报告来自同一隔离实验且读取不追加调用__254004
assert before["calibration"]["receipt_id"] is None assert before["calibration"]["receipt_id"] is None
assert after["calibration"]["receipt_id"] assert after["calibration"]["receipt_id"]
assert before["samples"] == after["samples"] assert before["samples"] == after["samples"]
elif report_kind == "effect":
assert before["effect"]["receipt_id"] is None and after["effect"]["receipt_id"]
assert before["effect"]["assessment"] == after["effect"]["assessment"]
assert before["samples"] == after["samples"]
else: else:
assert after == before assert after == before
assert len(env["received"]) == calls_before assert len(env["received"]) == calls_before