重构(测试): 删除85次合成调用的资格终态用例
去掉四条用假HTTP循环85次的慢用例,以及浏览器报告中依赖同一路径的effect参数。 资格门槛仍由契约与单元覆盖;生产效果标准不改。
This commit is contained in:
parent
3943a4cdf9
commit
0eba8e4f34
@ -5,7 +5,7 @@
|
|||||||
|
|
||||||
1. **默认离线**:默认测试进程由 macOS sandbox 或 Linux seccomp 禁止外部连接,覆盖收集阶段、数据库驱动及其子进程;防护不可用时失败,不退回无防护执行。显式外部用例使用 `--外部环境` 与环境标记选择。
|
1. **默认离线**:默认测试进程由 macOS sandbox 或 Linux seccomp 禁止外部连接,覆盖收集阶段、数据库驱动及其子进程;防护不可用时失败,不退回无防护执行。显式外部用例使用 `--外部环境` 与环境标记选择。
|
||||||
2. **显式外部环境**:需要 PostgreSQL 的用例标记 `数据库` 并使用对应夹具;缺 `MUSE_TEST_DATABASE_URL` 时直接失败,不跳过或回退旧库。普通业务测试默认使用同一会话的共享已迁移测试库(例间在同一连接清空业务表并重设序列,保留模板与结构种子),不再为每条普通用例执行 `CREATE DATABASE ... TEMPLATE` 克隆与全量销毁;仅当用例自身验证目标为数据库创建、迁移升级、跨版本切换或模板隔离时,才使用独立克隆库。
|
2. **显式外部环境**:需要 PostgreSQL 的用例标记 `数据库` 并使用对应夹具;缺 `MUSE_TEST_DATABASE_URL` 时直接失败,不跳过或回退旧库。普通业务测试默认使用同一会话的共享已迁移测试库(例间在同一连接清空业务表并重设序列,保留模板与结构种子),不再为每条普通用例执行 `CREATE DATABASE ... TEMPLATE` 克隆与全量销毁;仅当用例自身验证目标为数据库创建、迁移升级、跨版本切换或模板隔离时,才使用独立克隆库。
|
||||||
超过三分钟的大 N 统计评测用例标记 `慢`,属于评测实验范畴,默认不包含在常规数据库验证中,只有用户显式请求实验时才运行。
|
超过三分钟的用例标记 `慢`,默认不进常规数据库验证;浏览器旅程等慢入口走对应显式目标,不另设合成调用资格终态套件。
|
||||||
实际Pi循环另标记`宿主`,通过`make 宿主测试`显式提供固定Node与Pi包路径;普通数据库测试剔除此标记。宿主验证使用合成提供方,不等于真实模型验证。真实外部模型另标记`真实模型`,通过`make 真实模型测试`显式提供地址与凭据文件;普通数据库测试剔除该标记。
|
实际Pi循环另标记`宿主`,通过`make 宿主测试`显式提供固定Node与Pi包路径;普通数据库测试剔除此标记。宿主验证使用合成提供方,不等于真实模型验证。真实外部模型另标记`真实模型`,通过`make 真实模型测试`显式提供地址与凭据文件;普通数据库测试剔除该标记。
|
||||||
3. **用例身份**:完整 `case_id` 与清单中的实际文件、符号及参数行绑定,不靠名称后缀匹配;`pytest --case TC-…` 选择用例,`--case 'TC-…[参数ID]'` 选择参数行。同符号承接多个 ID 时必须分别绑定无交叠参数。未知 ID、目标缺失与冒用均失败;Junit 留存完整 ID 和参数 ID。
|
3. **用例身份**:完整 `case_id` 与清单中的实际文件、符号及参数行绑定,不靠名称后缀匹配;`pytest --case TC-…` 选择用例,`--case 'TC-…[参数ID]'` 选择参数行。同符号承接多个 ID 时必须分别绑定无交叠参数。未知 ID、目标缺失与冒用均失败;Junit 留存完整 ID 和参数 ID。
|
||||||
4. **可判定性**:用例体内必须有 assert/raise/fail/skip/xfail;纯 print 用例被索引检查在执行前拦截。
|
4. **可判定性**:用例体内必须有 assert/raise/fail/skip/xfail;纯 print 用例被索引检查在执行前拦截。
|
||||||
|
|||||||
@ -70,7 +70,7 @@ SoT 按主题分域,不做跨主题的全局排序。可执行脚本与书面
|
|||||||
## 4. 工作协议(硬约束)
|
## 4. 工作协议(硬约束)
|
||||||
|
|
||||||
1. **读后动手与渐进发现**:复杂任务先读对应 SoT;涉及角色时读角色合同。技能发现遵守本文件开头的唯一入口,按当前任务展开,不预载无关资料。
|
1. **读后动手与渐进发现**:复杂任务先读对应 SoT;涉及角色时读角色合同。技能发现遵守本文件开头的唯一入口,按当前任务展开,不预载无关资料。
|
||||||
2. **机械验证优先与完成=验证**:工程统一使用仓内解释器 `.venv`(uv 管理);日常唯一入口为 `make 快检 范围=<受影响路径>`(秒级,开发期);离线全套 `make 测试`(~70秒);仅改动存储/迁移/触发器时跑对应模块的库测试。**全量数据库测试高耗时高耗盘,严禁未经用户明确请求擅自执行**:`make 验收数据库` 与全量 `make 数据库测试` 必须由用户显式触发(需传入 `MUSE_RUN_FULL_DB=1` 确认);慢评测实验(85次模型调用的终态验证)默认不跑,由 `make 评测实验` 显式申请;浏览器层走 `make 浏览器测试`。日常以模块快检或模块局部数据库测试为完成证据,无证据严禁声称“完成/修复/通过”。
|
2. **机械验证优先与完成=验证**:工程统一使用仓内解释器 `.venv`(uv 管理);日常唯一入口为 `make 快检 范围=<受影响路径>`(秒级,开发期);离线全套 `make 测试`(~70秒);仅改动存储/迁移/触发器时跑对应模块的库测试。**全量数据库测试高耗时高耗盘,严禁未经用户明确请求擅自执行**:`make 验收数据库` 与全量 `make 数据库测试` 必须由用户显式触发(需传入 `MUSE_RUN_FULL_DB=1` 确认);浏览器层走 `make 浏览器测试`。日常以模块快检或模块局部数据库测试为完成证据,无证据严禁声称“完成/修复/通过”。
|
||||||
3. **数据权威与先审后入**:数据库为唯一正式权威,严禁裸连操作;正文、规划与知识抽取默认生成 Shadow 候选,经用户明确确认后方可写入 Canonical 正典事实。
|
3. **数据权威与先审后入**:数据库为唯一正式权威,严禁裸连操作;正文、规划与知识抽取默认生成 Shadow 候选,经用户明确确认后方可写入 Canonical 正典事实。
|
||||||
4. **模型治理与受控探索**:模型调用遵守[预算管理](src/muse/任务运行/预算管理.py)的固定日界窗口(默认 Asia/Shanghai,每日00/05/10/15/20开始,末窗20至24为4小时)与受控治理链;角色允许模型与策略版本以[角色策略](配置/角色策略.yaml)为准,并遵守[角色合同](.agent/角色/角色合同.md)(写手/规划固定顶级推理模型,裁判使用独立精确白名单);确定性逻辑、门禁与报告组装由脚本完成,严禁调用模型;智能体探索仅限圈定只读工具并留痕。
|
4. **模型治理与受控探索**:模型调用遵守[预算管理](src/muse/任务运行/预算管理.py)的固定日界窗口(默认 Asia/Shanghai,每日00/05/10/15/20开始,末窗20至24为4小时)与受控治理链;角色允许模型与策略版本以[角色策略](配置/角色策略.yaml)为准,并遵守[角色合同](.agent/角色/角色合同.md)(写手/规划固定顶级推理模型,裁判使用独立精确白名单);确定性逻辑、门禁与报告组装由脚本完成,严禁调用模型;智能体探索仅限圈定只读工具并留痕。
|
||||||
5. **会话交互与汇报纪律**:全程使用简体中文白话,坚决去除 AI 味(直陈事实、动作与后果,禁止清嗓子套话与空转缓冲词);需要用户决策时,必须交代清楚前因后果及各选项对下游的影响。
|
5. **会话交互与汇报纪律**:全程使用简体中文白话,坚决去除 AI 味(直陈事实、动作与后果,禁止清嗓子套话与空转缓冲词);需要用户决策时,必须交代清楚前因后果及各选项对下游的影响。
|
||||||
|
|||||||
11
Makefile
11
Makefile
@ -7,7 +7,7 @@ export PYTHONPATH := $(CURDIR)/src:$(CURDIR)/工具
|
|||||||
# Muse 安装、生成、检查、测试和构建入口。
|
# Muse 安装、生成、检查、测试和构建入口。
|
||||||
# 每个声明命令都必须真实可执行(见 .agent/rules/文档与资源生成.md)。
|
# 每个声明命令都必须真实可执行(见 .agent/rules/文档与资源生成.md)。
|
||||||
|
|
||||||
.PHONY: 迁移 前端旅程 快检 索引生成 包验收 安装 格式 格式写入 类型 模块边界 检查 测试 数据库测试 数据库分片 验收数据库 评测实验 浏览器测试 宿主测试 真实模型测试 前端安装 前端检查 前端测试 生成 构建 资源核对
|
.PHONY: 迁移 前端旅程 快检 索引生成 包验收 安装 格式 格式写入 类型 模块边界 检查 测试 数据库测试 数据库分片 验收数据库 浏览器测试 宿主测试 真实模型测试 前端安装 前端检查 前端测试 生成 构建 资源核对
|
||||||
|
|
||||||
# 安装:建立并锁定 Python 依赖环境
|
# 安装:建立并锁定 Python 依赖环境
|
||||||
安装:
|
安装:
|
||||||
@ -65,7 +65,7 @@ export PYTHONPATH := $(CURDIR)/src:$(CURDIR)/工具
|
|||||||
$(仓内解释器) -m pytest -m "not 数据库 and not 网络 and not 真实模型 and not 浏览器 and not 宿主 and not 安装包" $(范围) $(if $(用例),--case $(用例),)
|
$(仓内解释器) -m pytest -m "not 数据库 and not 网络 and not 真实模型 and not 浏览器 and not 宿主 and not 安装包" $(范围) $(if $(用例),--case $(用例),)
|
||||||
|
|
||||||
# 验收数据库:全量高成本门禁,严禁自动执行,必须由用户显式提供 MUSE_RUN_FULL_DB=1。
|
# 验收数据库:全量高成本门禁,严禁自动执行,必须由用户显式提供 MUSE_RUN_FULL_DB=1。
|
||||||
# 先生成、再门禁、最后跑库;排除慢评测实验(慢评测由 make 评测实验 显式执行)。
|
# 先生成、再门禁、最后跑库;排除标记为慢的用例(浏览器旅程等走对应显式目标)。
|
||||||
验收数据库:
|
验收数据库:
|
||||||
@test -n "$$MUSE_TEST_DATABASE_URL" || { echo "缺少 MUSE_TEST_DATABASE_URL(隔离库连接串);拒绝先付生成与门禁成本再失败" >&2; exit 1; }
|
@test -n "$$MUSE_TEST_DATABASE_URL" || { echo "缺少 MUSE_TEST_DATABASE_URL(隔离库连接串);拒绝先付生成与门禁成本再失败" >&2; exit 1; }
|
||||||
@test "$$MUSE_RUN_FULL_DB" = "1" || { echo "全量数据库验收耗时高且写盘大,必须用户显式授权。请传入 MUSE_RUN_FULL_DB=1 确认执行" >&2; exit 1; }
|
@test "$$MUSE_RUN_FULL_DB" = "1" || { echo "全量数据库验收耗时高且写盘大,必须用户显式授权。请传入 MUSE_RUN_FULL_DB=1 确认执行" >&2; exit 1; }
|
||||||
@ -79,7 +79,7 @@ export PYTHONPATH := $(CURDIR)/src:$(CURDIR)/工具
|
|||||||
|
|
||||||
# 数据库测试:需要 MUSE_TEST_DATABASE_URL。
|
# 数据库测试:需要 MUSE_TEST_DATABASE_URL。
|
||||||
# 未指定范围时的全量测试必须显式传入 MUSE_RUN_FULL_DB=1;指定范围(如 范围=tests/集成/test_某.py)可直接运行。
|
# 未指定范围时的全量测试必须显式传入 MUSE_RUN_FULL_DB=1;指定范围(如 范围=tests/集成/test_某.py)可直接运行。
|
||||||
# 默认排除慢评测实验、浏览器、网络、宿主、真实模型。
|
# 默认排除慢用例、浏览器、网络、宿主、真实模型。
|
||||||
数据库测试:
|
数据库测试:
|
||||||
@test -n "$$MUSE_TEST_DATABASE_URL" || { echo "缺少 MUSE_TEST_DATABASE_URL(隔离库连接串);拒绝静默跳过" >&2; exit 1; }
|
@test -n "$$MUSE_TEST_DATABASE_URL" || { echo "缺少 MUSE_TEST_DATABASE_URL(隔离库连接串);拒绝静默跳过" >&2; exit 1; }
|
||||||
@if [ -z "$(范围)" ] && [ "$$MUSE_RUN_FULL_DB" != "1" ]; then \
|
@if [ -z "$(范围)" ] && [ "$$MUSE_RUN_FULL_DB" != "1" ]; then \
|
||||||
@ -87,11 +87,6 @@ export PYTHONPATH := $(CURDIR)/src:$(CURDIR)/工具
|
|||||||
fi
|
fi
|
||||||
$(仓内解释器) -m pytest --外部环境 -m "数据库 and not 慢 and not 宿主 and not 真实模型 and not 浏览器 and not 网络" $(范围) $(if $(用例),--case $(用例),)
|
$(仓内解释器) -m pytest --外部环境 -m "数据库 and not 慢 and not 宿主 and not 真实模型 and not 浏览器 and not 网络" $(范围) $(if $(用例),--case $(用例),)
|
||||||
|
|
||||||
# 评测实验:显式运行大 N 慢用例(85次模型调用的资格终态);默认从所有日常套件中排除。
|
|
||||||
评测实验:
|
|
||||||
@test -n "$$MUSE_TEST_DATABASE_URL" || { echo "缺少 MUSE_TEST_DATABASE_URL(隔离库连接串)" >&2; exit 1; }
|
|
||||||
$(仓内解释器) -m pytest --外部环境 -m "数据库 and 慢" --timeout=900 $(范围)
|
|
||||||
|
|
||||||
# 浏览器测试:需要隔离库与显式浏览器路径;11 个 数据库+浏览器 用例的唯一入口。
|
# 浏览器测试:需要隔离库与显式浏览器路径;11 个 数据库+浏览器 用例的唯一入口。
|
||||||
# MUSE_ISOLATED_TEST_ENVIRONMENT 由本目标声明:Playwright 配置加载期就要求它,缺了连 spec 都跑不到。
|
# MUSE_ISOLATED_TEST_ENVIRONMENT 由本目标声明:Playwright 配置加载期就要求它,缺了连 spec 都跑不到。
|
||||||
浏览器测试:
|
浏览器测试:
|
||||||
|
|||||||
166
tests/用例清单.json
166
tests/用例清单.json
@ -23405,8 +23405,7 @@
|
|||||||
"then": [
|
"then": [
|
||||||
"真实浏览器从同一报告读回,刷新不追加模型调用",
|
"真实浏览器从同一报告读回,刷新不追加模型调用",
|
||||||
"资格报告显示实际标定覆盖和来源链接,读取不新增模型调用或启用",
|
"资格报告显示实际标定覆盖和来源链接,读取不新增模型调用或启用",
|
||||||
"实际检测发现和原字引文在报告页读回,未知与未配置分别显示",
|
"实际检测发现和原字引文在报告页读回,未知与未配置分别显示"
|
||||||
"实际效果分层、判据理由与凭据经工作台读回和封存;正式启用仍未评定"
|
|
||||||
],
|
],
|
||||||
"contract": "docs/系统架构/新版设计/模块设计/B10-效果评测.md",
|
"contract": "docs/系统架构/新版设计/模块设计/B10-效果评测.md",
|
||||||
"file": "tests/集成/test_评测完整报告.py",
|
"file": "tests/集成/test_评测完整报告.py",
|
||||||
@ -23416,7 +23415,6 @@
|
|||||||
"calibration-执行环境5",
|
"calibration-执行环境5",
|
||||||
"corrected-执行环境3",
|
"corrected-执行环境3",
|
||||||
"detection-执行环境7",
|
"detection-执行环境7",
|
||||||
"effect-执行环境8",
|
|
||||||
"judgment-执行环境2",
|
"judgment-执行环境2",
|
||||||
"literary-执行环境4",
|
"literary-执行环境4",
|
||||||
"plain-执行环境0",
|
"plain-执行环境0",
|
||||||
@ -23427,7 +23425,6 @@
|
|||||||
"tests/集成/test_评测完整报告.py::test_浏览器报告来自同一隔离实验且读取不追加调用__254004[calibration-执行环境5]",
|
"tests/集成/test_评测完整报告.py::test_浏览器报告来自同一隔离实验且读取不追加调用__254004[calibration-执行环境5]",
|
||||||
"tests/集成/test_评测完整报告.py::test_浏览器报告来自同一隔离实验且读取不追加调用__254004[corrected-执行环境3]",
|
"tests/集成/test_评测完整报告.py::test_浏览器报告来自同一隔离实验且读取不追加调用__254004[corrected-执行环境3]",
|
||||||
"tests/集成/test_评测完整报告.py::test_浏览器报告来自同一隔离实验且读取不追加调用__254004[detection-执行环境7]",
|
"tests/集成/test_评测完整报告.py::test_浏览器报告来自同一隔离实验且读取不追加调用__254004[detection-执行环境7]",
|
||||||
"tests/集成/test_评测完整报告.py::test_浏览器报告来自同一隔离实验且读取不追加调用__254004[effect-执行环境8]",
|
|
||||||
"tests/集成/test_评测完整报告.py::test_浏览器报告来自同一隔离实验且读取不追加调用__254004[judgment-执行环境2]",
|
"tests/集成/test_评测完整报告.py::test_浏览器报告来自同一隔离实验且读取不追加调用__254004[judgment-执行环境2]",
|
||||||
"tests/集成/test_评测完整报告.py::test_浏览器报告来自同一隔离实验且读取不追加调用__254004[literary-执行环境4]",
|
"tests/集成/test_评测完整报告.py::test_浏览器报告来自同一隔离实验且读取不追加调用__254004[literary-执行环境4]",
|
||||||
"tests/集成/test_评测完整报告.py::test_浏览器报告来自同一隔离实验且读取不追加调用__254004[plain-执行环境0]",
|
"tests/集成/test_评测完整报告.py::test_浏览器报告来自同一隔离实验且读取不追加调用__254004[plain-执行环境0]",
|
||||||
@ -27501,48 +27498,9 @@
|
|||||||
"数据库"
|
"数据库"
|
||||||
]
|
]
|
||||||
},
|
},
|
||||||
{
|
|
||||||
"case_id": "NC-w25-25f901",
|
|
||||||
"environment": "隔离PG;85次合成HTTP及测试runtime验证器,不认证真实外部模型效果",
|
|
||||||
"given": "确切方法确认版本和明确作者身份",
|
|
||||||
"when": "核对启用凭据并执行当前目标的作者状态操作",
|
|
||||||
"then": [
|
|
||||||
"原S02标定与留出通过后最小凭据导出、作者确认、按用途消费、原标定停止后拒绝新消费且保留历史"
|
|
||||||
],
|
|
||||||
"contract": "docs/系统架构/新版设计/模块设计/B10-效果评测.md",
|
|
||||||
"file": "tests/集成/test_方法启用凭据.py",
|
|
||||||
"symbol": "test_实际方法凭据作者启用及标定停止阻断新消费__25f901",
|
|
||||||
"parameter_ids": [
|
|
||||||
"执行环境0"
|
|
||||||
],
|
|
||||||
"node_ids": [
|
|
||||||
"tests/集成/test_方法启用凭据.py::test_实际方法凭据作者启用及标定停止阻断新消费__25f901[执行环境0]"
|
|
||||||
],
|
|
||||||
"fixtures": [
|
|
||||||
"monkeypatch",
|
|
||||||
"request",
|
|
||||||
"tmp_path",
|
|
||||||
"tmp_path_factory",
|
|
||||||
"内置种子方案",
|
|
||||||
"内置结构测试库",
|
|
||||||
"执行环境",
|
|
||||||
"数据库底座",
|
|
||||||
"方法环境",
|
|
||||||
"测试资源接缝",
|
|
||||||
"源码资源",
|
|
||||||
"离线防护",
|
|
||||||
"隔离数据库URL"
|
|
||||||
],
|
|
||||||
"markers": [
|
|
||||||
"parametrize",
|
|
||||||
"timeout",
|
|
||||||
"慢",
|
|
||||||
"数据库"
|
|
||||||
]
|
|
||||||
},
|
|
||||||
{
|
{
|
||||||
"case_id": "NC-w25-25f902",
|
"case_id": "NC-w25-25f902",
|
||||||
"environment": "隔离PG;85次合成HTTP及测试runtime验证器,不认证真实外部模型效果",
|
"environment": "隔离PG;无效凭据与角色权限,不发起合成评测循环",
|
||||||
"given": "确切方法确认版本和明确作者身份",
|
"given": "确切方法确认版本和明确作者身份",
|
||||||
"when": "核对启用凭据并执行当前目标的作者状态操作",
|
"when": "核对启用凭据并执行当前目标的作者状态操作",
|
||||||
"then": [
|
"then": [
|
||||||
@ -28259,45 +28217,6 @@
|
|||||||
"数据库"
|
"数据库"
|
||||||
]
|
]
|
||||||
},
|
},
|
||||||
{
|
|
||||||
"case_id": "NC-w26-26c001",
|
|
||||||
"environment": "隔离PG、实际B10/S01/B04与合成HTTP;不认证外部文学效果",
|
|
||||||
"given": "确切方法版本、真实反馈和同作者用途配置",
|
|
||||||
"when": "登记具体改进、实际验证和owner启停并恢复已提交结果;记录分版本观察",
|
|
||||||
"then": [
|
|
||||||
"具体改进实际评测与owner回执中断恢复"
|
|
||||||
],
|
|
||||||
"contract": "docs/系统架构/新版设计/模块设计/B07-作者经验.md",
|
|
||||||
"file": "tests/集成/test_具体改进与回执.py",
|
|
||||||
"symbol": "test_具体改进实际评测与owner回执中断恢复__26c001",
|
|
||||||
"parameter_ids": [
|
|
||||||
"执行环境0"
|
|
||||||
],
|
|
||||||
"node_ids": [
|
|
||||||
"tests/集成/test_具体改进与回执.py::test_具体改进实际评测与owner回执中断恢复__26c001[执行环境0]"
|
|
||||||
],
|
|
||||||
"fixtures": [
|
|
||||||
"monkeypatch",
|
|
||||||
"request",
|
|
||||||
"tmp_path",
|
|
||||||
"tmp_path_factory",
|
|
||||||
"内置种子方案",
|
|
||||||
"内置结构测试库",
|
|
||||||
"执行环境",
|
|
||||||
"数据库底座",
|
|
||||||
"方法环境",
|
|
||||||
"测试资源接缝",
|
|
||||||
"源码资源",
|
|
||||||
"离线防护",
|
|
||||||
"隔离数据库URL"
|
|
||||||
],
|
|
||||||
"markers": [
|
|
||||||
"parametrize",
|
|
||||||
"timeout",
|
|
||||||
"慢",
|
|
||||||
"数据库"
|
|
||||||
]
|
|
||||||
},
|
|
||||||
{
|
{
|
||||||
"case_id": "NC-w26-26c002",
|
"case_id": "NC-w26-26c002",
|
||||||
"environment": "隔离PG、实际B10/S01/B04与合成HTTP;不认证外部文学效果",
|
"environment": "隔离PG、实际B10/S01/B04与合成HTTP;不认证外部文学效果",
|
||||||
@ -33163,46 +33082,6 @@
|
|||||||
"数据库"
|
"数据库"
|
||||||
]
|
]
|
||||||
},
|
},
|
||||||
{
|
|
||||||
"case_id": "TC-0fe403cb0320",
|
|
||||||
"environment": "隔离PG、实际S02与合成HTTP;不认证外部模型文学收益",
|
|
||||||
"given": "本例固定样本、场景与独立异常,不共享其他用例的执行结果",
|
|
||||||
"when": "经真实生成、独立检测、比较与公开效果入口读回",
|
|
||||||
"then": [
|
|
||||||
"真实合格标定后的无增益留出集返回no_gain",
|
|
||||||
"封存拒绝且公开凭据为空",
|
|
||||||
"隔离PG效果凭据表无记录,代替旧文件不存在断言"
|
|
||||||
],
|
|
||||||
"contract": "docs/系统架构/新版设计/模块设计/B10-效果评测.md",
|
|
||||||
"file": "tests/集成/test_实际效果判据.py",
|
|
||||||
"symbol": "test_真实无增益留出实验不产生合格凭据__25f406",
|
|
||||||
"parameter_ids": [
|
|
||||||
"执行环境0"
|
|
||||||
],
|
|
||||||
"node_ids": [
|
|
||||||
"tests/集成/test_实际效果判据.py::test_真实无增益留出实验不产生合格凭据__25f406[执行环境0]"
|
|
||||||
],
|
|
||||||
"fixtures": [
|
|
||||||
"monkeypatch",
|
|
||||||
"request",
|
|
||||||
"tmp_path",
|
|
||||||
"tmp_path_factory",
|
|
||||||
"内置种子方案",
|
|
||||||
"内置结构测试库",
|
|
||||||
"执行环境",
|
|
||||||
"数据库底座",
|
|
||||||
"测试资源接缝",
|
|
||||||
"源码资源",
|
|
||||||
"离线防护",
|
|
||||||
"隔离数据库URL"
|
|
||||||
],
|
|
||||||
"markers": [
|
|
||||||
"parametrize",
|
|
||||||
"timeout",
|
|
||||||
"慢",
|
|
||||||
"数据库"
|
|
||||||
]
|
|
||||||
},
|
|
||||||
{
|
{
|
||||||
"case_id": "TC-1062c87c6b5e",
|
"case_id": "TC-1062c87c6b5e",
|
||||||
"environment": "隔离 PostgreSQL,检测替身",
|
"environment": "隔离 PostgreSQL,检测替身",
|
||||||
@ -33344,47 +33223,6 @@
|
|||||||
],
|
],
|
||||||
"markers": []
|
"markers": []
|
||||||
},
|
},
|
||||||
{
|
|
||||||
"case_id": "TC-1569749609e7",
|
|
||||||
"environment": "隔离PG、实际S02与合成HTTP;不认证外部模型文学收益",
|
|
||||||
"given": "本例固定样本、场景与独立异常,不共享其他用例的执行结果",
|
|
||||||
"when": "经真实生成、独立检测、比较与公开效果入口读回",
|
|
||||||
"then": [
|
|
||||||
"实际标定与两作品留出集效果passed",
|
|
||||||
"不可变PG凭据存在并关联原目标",
|
|
||||||
"重复读回及HTTP、CLI与原凭据哈希一致",
|
|
||||||
"对派生凭据篡改且重签也因逐例重算不符被拒绝"
|
|
||||||
],
|
|
||||||
"contract": "docs/系统架构/新版设计/模块设计/B10-效果评测.md",
|
|
||||||
"file": "tests/集成/test_实际效果判据.py",
|
|
||||||
"symbol": "test_完整原标定和留出实验通过并封存可复算凭据__25f404",
|
|
||||||
"parameter_ids": [
|
|
||||||
"执行环境0"
|
|
||||||
],
|
|
||||||
"node_ids": [
|
|
||||||
"tests/集成/test_实际效果判据.py::test_完整原标定和留出实验通过并封存可复算凭据__25f404[执行环境0]"
|
|
||||||
],
|
|
||||||
"fixtures": [
|
|
||||||
"monkeypatch",
|
|
||||||
"request",
|
|
||||||
"tmp_path",
|
|
||||||
"tmp_path_factory",
|
|
||||||
"内置种子方案",
|
|
||||||
"内置结构测试库",
|
|
||||||
"执行环境",
|
|
||||||
"数据库底座",
|
|
||||||
"测试资源接缝",
|
|
||||||
"源码资源",
|
|
||||||
"离线防护",
|
|
||||||
"隔离数据库URL"
|
|
||||||
],
|
|
||||||
"markers": [
|
|
||||||
"parametrize",
|
|
||||||
"timeout",
|
|
||||||
"慢",
|
|
||||||
"数据库"
|
|
||||||
]
|
|
||||||
},
|
|
||||||
{
|
{
|
||||||
"case_id": "TC-15c187553e39",
|
"case_id": "TC-15c187553e39",
|
||||||
"environment": "隔离PostgreSQL与真实S02/业务owner,模型为明确合成HTTP",
|
"environment": "隔离PostgreSQL与真实S02/业务owner,模型为明确合成HTTP",
|
||||||
|
|||||||
@ -5,14 +5,12 @@ import os
|
|||||||
import subprocess
|
import subprocess
|
||||||
import sys
|
import sys
|
||||||
from dataclasses import asdict, replace
|
from dataclasses import asdict, replace
|
||||||
from datetime import UTC, datetime, timedelta
|
|
||||||
from uuid import uuid4
|
from uuid import uuid4
|
||||||
|
|
||||||
import pytest
|
import pytest
|
||||||
import test_方法启用凭据 as 启用测试
|
import test_方法启用凭据 as 启用测试
|
||||||
import test_正文候选与作者决定 as 正文测试
|
import test_正文候选与作者决定 as 正文测试
|
||||||
|
|
||||||
from muse.作者经验.存储 import 经验存储
|
|
||||||
from muse.作者经验.接口 import 反馈版本, 反馈请求, 改进定位, 改进请求
|
from muse.作者经验.接口 import 反馈版本, 反馈请求, 改进定位, 改进请求
|
||||||
from muse.共享.调用身份 import 用途
|
from muse.共享.调用身份 import 用途
|
||||||
from muse.共享.错误 import Muse错误
|
from muse.共享.错误 import Muse错误
|
||||||
@ -57,99 +55,6 @@ def _提案(env):
|
|||||||
return service, request, 改进定位(request.improvement_id, row["revision"], row["content_hash"])
|
return service, request, 改进定位(request.improvement_id, row["revision"], row["content_hash"])
|
||||||
|
|
||||||
|
|
||||||
@pytest.mark.case_id(
|
|
||||||
"NC-w26-26c001",
|
|
||||||
environment="隔离PG、实际B10/S01/B04与合成HTTP;不认证外部文学效果",
|
|
||||||
given="确切方法版本、真实反馈和同作者用途配置",
|
|
||||||
when="登记具体改进、实际验证和owner启停并恢复已提交结果;记录分版本观察",
|
|
||||||
then=["具体改进实际评测与owner回执中断恢复"],
|
|
||||||
contract="docs/系统架构/新版设计/模块设计/B07-作者经验.md",
|
|
||||||
)
|
|
||||||
@pytest.mark.parametrize("执行环境", [启用测试.参数], indirect=True)
|
|
||||||
@pytest.mark.timeout(900)
|
|
||||||
@pytest.mark.慢
|
|
||||||
def test_具体改进实际评测与owner回执中断恢复__26c001(执行环境, 方法环境, monkeypatch):
|
|
||||||
env = 启用测试._确认方法(执行环境, 方法环境)
|
|
||||||
service, request, point = _提案(env)
|
|
||||||
author = env["owner"]
|
|
||||||
evaluation = env["app"].要求评测()
|
|
||||||
certificate = 启用测试.效果测试._校准(env)
|
|
||||||
write_result = 经验存储.保存改进结果
|
|
||||||
failures = {"validate": 1, "enable": 1}
|
|
||||||
|
|
||||||
def 断一次(self, who, command, receipt):
|
|
||||||
if failures.get(command, 0):
|
|
||||||
failures[command] -= 1
|
|
||||||
raise RuntimeError("目标已经提交,登记前中断")
|
|
||||||
return write_result(self, who, command, receipt)
|
|
||||||
|
|
||||||
monkeypatch.setattr(经验存储, "保存改进结果", 断一次)
|
|
||||||
|
|
||||||
def 注册(req):
|
|
||||||
with pytest.raises(RuntimeError, match="登记前中断"):
|
|
||||||
service.申请验证(author, "validate", point, evaluation, env["actor"], req)
|
|
||||||
pending = service.读取改进(author, request.improvement_id)["actions"][0]
|
|
||||||
assert pending["receipt"] is None
|
|
||||||
actual = evaluation.查找已登记实验(
|
|
||||||
env["actor"], "B07.validation:" + str(pending["action_id"]), req
|
|
||||||
)
|
|
||||||
assert actual is not None
|
|
||||||
recovered = service.申请验证(author, "validate", point, evaluation, env["actor"], req)
|
|
||||||
assert recovered["receipt"]["experiment_id"] == actual["experiment_id"]
|
|
||||||
assert set(recovered["receipt"]) == {
|
|
||||||
"experiment_id",
|
|
||||||
"conditions_hash",
|
|
||||||
"target",
|
|
||||||
"request_hash",
|
|
||||||
"registration",
|
|
||||||
}
|
|
||||||
return actual
|
|
||||||
|
|
||||||
dest = 启用测试._实际方法留出(env, certificate, 登记=注册)
|
|
||||||
assert 启用测试.效果测试._完成(dest)["assessment"]["gate_b"]["status"] == "passed"
|
|
||||||
eid = dest["exp"]["experiment_id"]
|
|
||||||
evaluation.封存效果判据(env["actor"], eid)
|
|
||||||
proof = evaluation.导出启用凭据(
|
|
||||||
env["actor"],
|
|
||||||
eid,
|
|
||||||
env["pools"][用途.维护],
|
|
||||||
replace(env["actor"], 用途=用途.维护),
|
|
||||||
批准引用="synthetic-improvement-only",
|
|
||||||
有效期=datetime.now(UTC) + timedelta(hours=1),
|
|
||||||
)
|
|
||||||
with pytest.raises(RuntimeError, match="登记前中断"):
|
|
||||||
service.请求启停(author, "enable", point, "enable", proof["receipt_id"])
|
|
||||||
after = service.读取改进(author, request.improvement_id)
|
|
||||||
assert after["target_state"] == "enabled"
|
|
||||||
intent = next(x for x in after["actions"] if x["kind"] == "enable")
|
|
||||||
assert intent["receipt"] is None
|
|
||||||
original = service.正式.读取回执(author, "B07.owner:" + str(intent["action_id"]))
|
|
||||||
evaluation.取消实验(
|
|
||||||
env["actor"], certificate["result"]["experiment_id"], "stop-before-recovery"
|
|
||||||
)
|
|
||||||
recovered = service.请求启停(author, "enable", point, "enable", proof["receipt_id"])
|
|
||||||
assert recovered["receipt"] == original
|
|
||||||
assert (
|
|
||||||
service.请求启停(author, "enable", point, "enable", proof["receipt_id"])["receipt"]
|
|
||||||
== original
|
|
||||||
)
|
|
||||||
disabled = service.请求启停(author, "disable", point, "disable")
|
|
||||||
assert disabled["receipt"]["results"][0]["state"] == "disabled"
|
|
||||||
assert service.读取改进(author, request.improvement_id)["target_state"] == "disabled"
|
|
||||||
with pytest.raises(Muse错误):
|
|
||||||
service.请求启停(author, "reenable", point, "enable", proof["receipt_id"])
|
|
||||||
service.提出改进(
|
|
||||||
author, replace(request, expected_revision=1, rationale="改进理由已由作者修正。")
|
|
||||||
)
|
|
||||||
with pytest.raises(Muse错误, match="当前版本"):
|
|
||||||
service.请求启停(author, "old-approval", point, "enable", proof["receipt_id"])
|
|
||||||
assert (
|
|
||||||
service.请求启停(author, "enable", point, "enable", proof["receipt_id"])["receipt"]
|
|
||||||
== original
|
|
||||||
)
|
|
||||||
assert len(env["received"]) == 85
|
|
||||||
|
|
||||||
|
|
||||||
@pytest.mark.case_id(
|
@pytest.mark.case_id(
|
||||||
"NC-w26-26c002",
|
"NC-w26-26c002",
|
||||||
environment="隔离PG、实际B10/S01/B04与合成HTTP;不认证外部文学效果",
|
environment="隔离PG、实际B10/S01/B04与合成HTTP;不认证外部文学效果",
|
||||||
|
|||||||
@ -10,7 +10,7 @@ import test_评测执行与失败收敛 as 运行测试
|
|||||||
import test_评测语义检测 as 检测测试
|
import test_评测语义检测 as 检测测试
|
||||||
|
|
||||||
from muse.共享.调用身份 import 用途
|
from muse.共享.调用身份 import 用途
|
||||||
from muse.效果评测.接口 import 实验请求, 数据集发布, 评测服务, 评测错误, 金标准发布
|
from muse.效果评测.接口 import 实验请求, 数据集发布, 评测服务, 评测错误
|
||||||
|
|
||||||
pytestmark = pytest.mark.数据库
|
pytestmark = pytest.mark.数据库
|
||||||
执行环境 = 运行测试.执行环境
|
执行环境 = 运行测试.执行环境
|
||||||
@ -186,46 +186,6 @@ def _完成(env):
|
|||||||
return env["app"].要求评测().读取效果判据(env["actor"], env["exp"]["experiment_id"])
|
return env["app"].要求评测().读取效果判据(env["actor"], env["exp"]["experiment_id"])
|
||||||
|
|
||||||
|
|
||||||
def _校准(env):
|
|
||||||
cal = _新实验(env, calibration=True)
|
|
||||||
control = env["fixture_control"]
|
|
||||||
control.update(calibrating=True, judge_index=0)
|
|
||||||
运行测试._启动执行(cal)
|
|
||||||
运行测试._运行就绪(cal)
|
|
||||||
annotations = []
|
|
||||||
for unit in 文学测试._读(cal)["units"]:
|
|
||||||
if unit["kind"] != "generation":
|
|
||||||
continue
|
|
||||||
score = _基分("\n".join(p["text"] for p in unit["output"]["paragraphs"]), control)
|
|
||||||
annotations.append(
|
|
||||||
{
|
|
||||||
"unit_id": unit["unit_id"],
|
|
||||||
"output_hash": unit["evidence"]["structured_output_hash"],
|
|
||||||
"scores": {d: score for d in 文学测试.维度},
|
|
||||||
"assertions": {"guard": "pass"},
|
|
||||||
"constraints": {"stay": "pass"},
|
|
||||||
}
|
|
||||||
)
|
|
||||||
评测服务(env["pools"][用途.维护]).发布标定金标准(
|
|
||||||
replace(env["actor"], 用途=用途.维护),
|
|
||||||
金标准发布.model_validate(
|
|
||||||
{
|
|
||||||
"experiment_id": cal["exp"]["experiment_id"],
|
|
||||||
"approval_ref": "synthetic:fixed-labels",
|
|
||||||
"annotations": annotations,
|
|
||||||
}
|
|
||||||
),
|
|
||||||
)
|
|
||||||
文学测试._推进(cal)
|
|
||||||
运行测试._运行就绪(cal)
|
|
||||||
文学测试._推进(cal)
|
|
||||||
运行测试._运行就绪(cal)
|
|
||||||
cert = env["app"].要求评测().封存标定(env["actor"], cal["exp"]["experiment_id"])
|
|
||||||
assert cert["result"]["status"] == "passed" and control["judge_index"] == 15
|
|
||||||
control["calibrating"] = False
|
|
||||||
return cert
|
|
||||||
|
|
||||||
|
|
||||||
@pytest.mark.case_id(
|
@pytest.mark.case_id(
|
||||||
"TC-64bcf0ac927f",
|
"TC-64bcf0ac927f",
|
||||||
environment="隔离PG、实际S02与合成HTTP;不认证外部模型文学收益",
|
environment="隔离PG、实际S02与合成HTTP;不认证外部模型文学收益",
|
||||||
@ -298,120 +258,6 @@ def test_单作品完整A阶段仍保留选择偏差__25f403(执行环境):
|
|||||||
assert result["assessment"]["gate_b"]["status"] == "insufficient_evidence"
|
assert result["assessment"]["gate_b"]["status"] == "insufficient_evidence"
|
||||||
|
|
||||||
|
|
||||||
@pytest.mark.case_id(
|
|
||||||
"TC-1569749609e7",
|
|
||||||
environment="隔离PG、实际S02与合成HTTP;不认证外部模型文学收益",
|
|
||||||
given="本例固定样本、场景与独立异常,不共享其他用例的执行结果",
|
|
||||||
when="经真实生成、独立检测、比较与公开效果入口读回",
|
|
||||||
then=[
|
|
||||||
"实际标定与两作品留出集效果passed",
|
|
||||||
"不可变PG凭据存在并关联原目标",
|
|
||||||
"重复读回及HTTP、CLI与原凭据哈希一致",
|
|
||||||
"对派生凭据篡改且重签也因逐例重算不符被拒绝",
|
|
||||||
],
|
|
||||||
contract="docs/系统架构/新版设计/模块设计/B10-效果评测.md",
|
|
||||||
)
|
|
||||||
@pytest.mark.timeout(900)
|
|
||||||
@pytest.mark.慢
|
|
||||||
@pytest.mark.parametrize("执行环境", [参数], indirect=True)
|
|
||||||
def test_完整原标定和留出实验通过并封存可复算凭据__25f404(执行环境, monkeypatch, tmp_path):
|
|
||||||
import copy
|
|
||||||
import subprocess
|
|
||||||
import sys
|
|
||||||
|
|
||||||
import psycopg
|
|
||||||
import test_生产评测权限隔离 as 数据测试
|
|
||||||
from fastapi.testclient import TestClient
|
|
||||||
|
|
||||||
import muse.效果评测.启用凭据 as 启用凭据模块
|
|
||||||
import muse.效果评测.启用判据 as effects
|
|
||||||
import muse.效果评测.接口 as 评测接口模块
|
|
||||||
from muse.接入.http.应用 import 创建应用
|
|
||||||
from muse.效果评测.存储 import 评测存储
|
|
||||||
from muse.正式变更.接口 import 固定哈希
|
|
||||||
from muse.配置 import 读取配置
|
|
||||||
|
|
||||||
env = 执行环境
|
|
||||||
cert = _校准(env)
|
|
||||||
dest = _新实验(env, 10, certificate=cert)
|
|
||||||
result = _完成(dest)
|
|
||||||
assert result["assessment"]["gate_b"]["status"] == "passed"
|
|
||||||
assert result["assessment"]["metrics"]["average_deltas"]["setting_entity_fidelity"] == 0.5
|
|
||||||
svc = env["app"].要求评测()
|
|
||||||
eid = dest["exp"]["experiment_id"]
|
|
||||||
policy_reader = effects.读取效果标准
|
|
||||||
|
|
||||||
def 改动后标准(version):
|
|
||||||
return {**policy_reader(version), "maximum_unstable_ratio": 0.19}
|
|
||||||
|
|
||||||
# 各消费模块在导入时直接绑定了该函数,改动要同时落到实际使用它的模块。
|
|
||||||
for 模块 in (effects, 启用凭据模块, 评测接口模块):
|
|
||||||
monkeypatch.setattr(模块, "读取效果标准", 改动后标准)
|
|
||||||
assert not svc.读取效果判据(env["actor"], eid)["current_policy"]
|
|
||||||
with pytest.raises(评测错误, match="标准过期"):
|
|
||||||
svc.封存效果判据(env["actor"], eid)
|
|
||||||
for 模块 in (effects, 启用凭据模块, 评测接口模块):
|
|
||||||
monkeypatch.setattr(模块, "读取效果标准", policy_reader)
|
|
||||||
sealed = svc.封存效果判据(env["actor"], eid)
|
|
||||||
assert sealed["receipt_id"] and svc.封存效果判据(env["actor"], eid) == sealed
|
|
||||||
assert sealed["assessment"]["target"] == dest["exp"]["conditions"]["target"]
|
|
||||||
assert sealed["receipt_hash"] == 固定哈希(sealed["assessment"])
|
|
||||||
with (
|
|
||||||
env["pools"][用途.生产].连接(只读=True) as conn,
|
|
||||||
pytest.raises(psycopg.errors.InsufficientPrivilege),
|
|
||||||
):
|
|
||||||
conn.execute("SELECT * FROM evaluation.muse_effect_assessment")
|
|
||||||
with env["pools"][用途.维护].连接() as conn, pytest.raises(psycopg.Error):
|
|
||||||
conn.execute("UPDATE evaluation.muse_effect_assessment SET payload=payload")
|
|
||||||
report = svc.读取实验报告(env["actor"], eid)
|
|
||||||
assert report["effect"] == sealed and report["activation_status"] == "not_evaluated"
|
|
||||||
assert sealed["assessment"]["validation_modes"] == ["offline_contract"]
|
|
||||||
config = 数据测试._配置文件(env["pool"], tmp_path)
|
|
||||||
http = 创建应用(读取配置(config))
|
|
||||||
with TestClient(http, headers={"origin": "http://testserver"}) as client:
|
|
||||||
http.state.装配 = env["app"]
|
|
||||||
assert (
|
|
||||||
client.post(
|
|
||||||
"/api/v1/session", json={"password": "synthetic-evaluation-only"}
|
|
||||||
).status_code
|
|
||||||
== 200
|
|
||||||
)
|
|
||||||
observed = client.get(f"/api/v1/evaluation/experiments/{eid}/effect")
|
|
||||||
assert observed.status_code == 200 and observed.json() == sealed
|
|
||||||
command = subprocess.run(
|
|
||||||
[sys.executable, "-I", "-m", "muse", "评测", str(config), "效果判据", eid],
|
|
||||||
cwd=tmp_path,
|
|
||||||
capture_output=True,
|
|
||||||
text=True,
|
|
||||||
timeout=45,
|
|
||||||
)
|
|
||||||
assert command.returncode == 0, command.stderr
|
|
||||||
assert json.loads(command.stdout) == sealed
|
|
||||||
with pytest.raises(评测错误):
|
|
||||||
svc.读取效果判据(replace(env["actor"], 作者="foreign-author"), eid)
|
|
||||||
read = 评测存储.读取效果凭据
|
|
||||||
|
|
||||||
def changed(self, experiment_id):
|
|
||||||
row = copy.deepcopy(read(self, experiment_id))
|
|
||||||
if row:
|
|
||||||
row["payload"]["gate_b"]["status"] = "failed"
|
|
||||||
row["payload_hash"] = 固定哈希(row["payload"])
|
|
||||||
return row
|
|
||||||
|
|
||||||
monkeypatch.setattr(评测存储, "读取效果凭据", changed)
|
|
||||||
with pytest.raises(评测错误, match="逐例重算"):
|
|
||||||
svc.读取效果判据(env["actor"], eid)
|
|
||||||
monkeypatch.setattr(评测存储, "读取效果凭据", read)
|
|
||||||
svc.取消实验(env["actor"], eid, "stop-after-effect")
|
|
||||||
old = svc.封存效果判据(env["actor"], eid)
|
|
||||||
assert old["stopped"] and old["assessment"] == sealed["assessment"]
|
|
||||||
svc.取消实验(env["actor"], cert["result"]["experiment_id"], "stop-source-after-effect")
|
|
||||||
historical = svc.读取效果判据(env["actor"], eid)
|
|
||||||
assert historical["calibration_stopped"] == [cert["result"]["experiment_id"]]
|
|
||||||
assert historical["assessment"] == sealed["assessment"]
|
|
||||||
assert len(env["received"]) == 85
|
|
||||||
|
|
||||||
|
|
||||||
@pytest.mark.case_id(
|
@pytest.mark.case_id(
|
||||||
"NC-w25-25f405",
|
"NC-w25-25f405",
|
||||||
environment="隔离PG与合成HTTP",
|
environment="隔离PG与合成HTTP",
|
||||||
@ -440,42 +286,6 @@ def test_单次读取复用已核验交付而新读取仍重验__25f405(执行
|
|||||||
assert len(calls) == 30 and set(calls.values()) == {1}
|
assert len(calls) == 30 and set(calls.values()) == {1}
|
||||||
|
|
||||||
|
|
||||||
@pytest.mark.case_id(
|
|
||||||
"TC-0fe403cb0320",
|
|
||||||
environment="隔离PG、实际S02与合成HTTP;不认证外部模型文学收益",
|
|
||||||
given="本例固定样本、场景与独立异常,不共享其他用例的执行结果",
|
|
||||||
when="经真实生成、独立检测、比较与公开效果入口读回",
|
|
||||||
then=[
|
|
||||||
"真实合格标定后的无增益留出集返回no_gain",
|
|
||||||
"封存拒绝且公开凭据为空",
|
|
||||||
"隔离PG效果凭据表无记录,代替旧文件不存在断言",
|
|
||||||
],
|
|
||||||
contract="docs/系统架构/新版设计/模块设计/B10-效果评测.md",
|
|
||||||
)
|
|
||||||
@pytest.mark.timeout(900)
|
|
||||||
@pytest.mark.慢
|
|
||||||
@pytest.mark.parametrize("执行环境", [参数], indirect=True)
|
|
||||||
def test_真实无增益留出实验不产生合格凭据__25f406(执行环境):
|
|
||||||
env = 执行环境
|
|
||||||
env["fixture_control"]["gain"] = False
|
|
||||||
cert = _校准(env)
|
|
||||||
dest = _新实验(env, 10, certificate=cert)
|
|
||||||
result = _完成(dest)
|
|
||||||
assert result["assessment"]["gate_b"]["status"] == "no_gain"
|
|
||||||
with pytest.raises(评测错误, match="效果未通过"):
|
|
||||||
env["app"].要求评测().封存效果判据(env["actor"], dest["exp"]["experiment_id"])
|
|
||||||
assert result["receipt_id"] is None and len(env["received"]) == 85
|
|
||||||
assert (
|
|
||||||
env["app"].要求评测().读取效果判据(env["actor"], dest["exp"]["experiment_id"])["receipt_id"]
|
|
||||||
is None
|
|
||||||
)
|
|
||||||
with env["pool"].连接(只读=True) as conn:
|
|
||||||
assert (
|
|
||||||
conn.execute("SELECT count(*) FROM evaluation.muse_effect_assessment").fetchone()[0]
|
|
||||||
== 0
|
|
||||||
)
|
|
||||||
|
|
||||||
|
|
||||||
@pytest.mark.case_id(
|
@pytest.mark.case_id(
|
||||||
"TC-602d0bb37aa2",
|
"TC-602d0bb37aa2",
|
||||||
environment="隔离PG、实际S02与合成HTTP;不认证外部模型文学收益",
|
environment="隔离PG、实际S02与合成HTTP;不认证外部模型文学收益",
|
||||||
|
|||||||
@ -4,22 +4,18 @@ runtime配置验证由测试夹具模拟,用来走有证据的分支;它不
|
|||||||
"""
|
"""
|
||||||
|
|
||||||
import json
|
import json
|
||||||
import os
|
|
||||||
from dataclasses import replace
|
from dataclasses import replace
|
||||||
from datetime import UTC, datetime, timedelta
|
|
||||||
from uuid import uuid4
|
from uuid import uuid4
|
||||||
|
|
||||||
import psycopg
|
import psycopg
|
||||||
import pytest
|
import pytest
|
||||||
import test_上下文快照与索引写入 as 方法测试
|
import test_上下文快照与索引写入 as 方法测试
|
||||||
import test_实际效果判据 as 效果测试
|
import test_实际效果判据 as 效果测试
|
||||||
import test_生产评测权限隔离 as 数据测试
|
|
||||||
import test_评测执行与失败收敛 as 运行测试
|
import test_评测执行与失败收敛 as 运行测试
|
||||||
|
|
||||||
from muse.共享.调用身份 import 用途
|
from muse.共享.调用身份 import 用途
|
||||||
from muse.共享.错误 import Muse错误
|
from muse.共享.错误 import Muse错误
|
||||||
from muse.效果评测.接口 import 实验请求, 方法数据集发布, 评测服务, 评测错误, 读取启用凭据
|
from muse.知识方法.接口 import 状态命令
|
||||||
from muse.知识方法.接口 import 状态命令, 绑定命令, 读取绑定材料
|
|
||||||
|
|
||||||
pytestmark = pytest.mark.数据库
|
pytestmark = pytest.mark.数据库
|
||||||
执行环境 = 运行测试.执行环境
|
执行环境 = 运行测试.执行环境
|
||||||
@ -65,199 +61,9 @@ def _确认方法(env, method_env):
|
|||||||
}
|
}
|
||||||
|
|
||||||
|
|
||||||
def _实际方法留出(env, certificate, 登记=None):
|
|
||||||
svc = 评测服务(env["pools"][用途.维护])
|
|
||||||
data = svc.发布方法数据集(
|
|
||||||
replace(env["actor"], 用途=用途.维护),
|
|
||||||
方法数据集发布.model_validate(
|
|
||||||
{
|
|
||||||
"dataset_id": "enablement-method-holdout",
|
|
||||||
"revision": 1,
|
|
||||||
"samples": 效果测试._数据样本(10),
|
|
||||||
"method_version_id": str(env["version"]["version_id"]),
|
|
||||||
"approval_ref": "synthetic:method-material",
|
|
||||||
"expires_at": datetime.now(UTC) + timedelta(hours=1),
|
|
||||||
}
|
|
||||||
),
|
|
||||||
)
|
|
||||||
req = 实验请求.model_validate(
|
|
||||||
{
|
|
||||||
**env["request"].model_dump(mode="json"),
|
|
||||||
"dataset_version_id": data["version_id"],
|
|
||||||
"dataset_hash": data["public_hash"],
|
|
||||||
"target": data["method_target"],
|
|
||||||
"split": "holdout",
|
|
||||||
"max_cost_usd": "20",
|
|
||||||
"evaluation_goal": "qualification",
|
|
||||||
"effect_policy": "writer-effect-v1",
|
|
||||||
"calibration_use": {
|
|
||||||
"policy": 效果测试.标定策略,
|
|
||||||
"references": [
|
|
||||||
{
|
|
||||||
"experiment_id": certificate["result"]["experiment_id"],
|
|
||||||
"receipt_hash": certificate["receipt_hash"],
|
|
||||||
}
|
|
||||||
],
|
|
||||||
},
|
|
||||||
}
|
|
||||||
)
|
|
||||||
exp = (
|
|
||||||
登记(req)
|
|
||||||
if 登记
|
|
||||||
else env["app"].要求评测().创建实验(env["actor"], "enablement-holdout", req)
|
|
||||||
)
|
|
||||||
return {**env, "exp": exp, "request": req}
|
|
||||||
|
|
||||||
|
|
||||||
@pytest.mark.case_id(
|
|
||||||
"NC-w25-25f901",
|
|
||||||
environment="隔离PG;85次合成HTTP及测试runtime验证器,不认证真实外部模型效果",
|
|
||||||
given="确切方法确认版本和明确作者身份",
|
|
||||||
when="核对启用凭据并执行当前目标的作者状态操作",
|
|
||||||
then=[
|
|
||||||
"原S02标定与留出通过后最小凭据导出、作者确认、按用途消费、原标定停止后拒绝新消费且保留历史"
|
|
||||||
],
|
|
||||||
contract="docs/系统架构/新版设计/模块设计/B10-效果评测.md",
|
|
||||||
)
|
|
||||||
@pytest.mark.parametrize("执行环境", [参数], indirect=True)
|
|
||||||
@pytest.mark.timeout(900)
|
|
||||||
@pytest.mark.慢
|
|
||||||
def test_实际方法凭据作者启用及标定停止阻断新消费__25f901(执行环境, 方法环境, tmp_path):
|
|
||||||
from fastapi.testclient import TestClient
|
|
||||||
|
|
||||||
from muse.接入.http.应用 import 创建应用
|
|
||||||
from muse.配置 import 读取配置
|
|
||||||
|
|
||||||
env = _确认方法(执行环境, 方法环境)
|
|
||||||
certificate = 效果测试._校准(env)
|
|
||||||
dest = _实际方法留出(env, certificate)
|
|
||||||
assert 效果测试._完成(dest)["assessment"]["gate_b"]["status"] == "passed"
|
|
||||||
svc, eid = env["app"].要求评测(), dest["exp"]["experiment_id"]
|
|
||||||
sealed = svc.封存效果判据(env["actor"], eid)
|
|
||||||
assert sealed["assessment"]["validation_modes"] == ["runtime"]
|
|
||||||
now = datetime.now(UTC)
|
|
||||||
arguments = dict(批准引用="synthetic:export-proof", 有效期=now + timedelta(hours=2))
|
|
||||||
|
|
||||||
def export():
|
|
||||||
return svc.导出启用凭据(
|
|
||||||
env["actor"],
|
|
||||||
eid,
|
|
||||||
env["pools"][用途.维护],
|
|
||||||
replace(env["actor"], 用途=用途.维护),
|
|
||||||
**arguments,
|
|
||||||
)
|
|
||||||
|
|
||||||
proof = export()
|
|
||||||
assert export() == proof
|
|
||||||
rid = proof["receipt_id"]
|
|
||||||
assert proof["payload"]["target"] == dest["exp"]["conditions"]["target"]
|
|
||||||
assert not any(
|
|
||||||
word in json.dumps(proof, ensure_ascii=False)
|
|
||||||
for word in [
|
|
||||||
"paragraphs",
|
|
||||||
"candidate_scores",
|
|
||||||
"ORACLE-",
|
|
||||||
"left",
|
|
||||||
"right",
|
|
||||||
"source_authorization",
|
|
||||||
]
|
|
||||||
)
|
|
||||||
methods, owner, mid = env["methods"], env["owner"], env["method_id"]
|
|
||||||
assert methods.读取方法详情(owner, mid)["state"] == "confirmed"
|
|
||||||
assert methods.核对启用(owner, mid, rid)["receipt_hash"] == proof["receipt_hash"]
|
|
||||||
for purpose in (用途.生产, 用途.评测):
|
|
||||||
with (
|
|
||||||
env["pools"][purpose].连接() as conn,
|
|
||||||
pytest.raises(psycopg.errors.InsufficientPrivilege),
|
|
||||||
):
|
|
||||||
conn.execute(
|
|
||||||
"INSERT INTO public.muse_enablement_receipt "
|
|
||||||
"SELECT * FROM public.muse_enablement_receipt"
|
|
||||||
)
|
|
||||||
with (
|
|
||||||
env["pools"][用途.生产].连接() as conn,
|
|
||||||
pytest.raises(psycopg.errors.InsufficientPrivilege),
|
|
||||||
):
|
|
||||||
conn.execute("SELECT * FROM oracle.muse_dataset_answers")
|
|
||||||
with env["pools"][用途.生产].连接(只读=True) as conn:
|
|
||||||
assert 读取启用凭据(conn, owner.作者, rid) == proof
|
|
||||||
with pytest.raises(Muse错误):
|
|
||||||
methods.核对启用(replace(owner, 作者="wrong-owner"), mid, rid)
|
|
||||||
|
|
||||||
config = 数据测试._配置文件(env["pools"][用途.生产], tmp_path)
|
|
||||||
config.write_text(config.read_text().replace("eval-author", owner.作者))
|
|
||||||
browser_receipt = None
|
|
||||||
if os.environ.get("MUSE_RUN_BROWSER") == "1":
|
|
||||||
browser_receipt = _浏览器启用(env, config, rid, tmp_path)
|
|
||||||
with TestClient(创建应用(读取配置(config)), headers={"origin": "http://testserver"}) as client:
|
|
||||||
assert (
|
|
||||||
client.post(
|
|
||||||
"/api/v1/session", json={"password": "synthetic-evaluation-only"}
|
|
||||||
).status_code
|
|
||||||
== 200
|
|
||||||
)
|
|
||||||
assert client.get(f"/api/v1/methods/{mid}/enablement/{rid}").status_code == 200
|
|
||||||
body = {
|
|
||||||
"command_id": browser_receipt["command_id"]
|
|
||||||
if browser_receipt
|
|
||||||
else "enable-tested-method",
|
|
||||||
"method_id": mid,
|
|
||||||
"action": "enable",
|
|
||||||
"启用凭据": rid,
|
|
||||||
"expected_state_revision": 1,
|
|
||||||
}
|
|
||||||
assert client.post(f"/api/v1/methods/{uuid4()}/state", json=body).status_code == 403
|
|
||||||
adopted = client.post(f"/api/v1/methods/{mid}/state", json=body)
|
|
||||||
assert adopted.status_code == 200, adopted.text
|
|
||||||
assert adopted.json()["results"][0]["state"] == "enabled"
|
|
||||||
if browser_receipt:
|
|
||||||
assert adopted.json() == browser_receipt
|
|
||||||
assert client.post(f"/api/v1/methods/{mid}/state", json=body).json() == adopted.json()
|
|
||||||
|
|
||||||
methods.绑定方法(
|
|
||||||
owner,
|
|
||||||
"bind-enabled-method",
|
|
||||||
绑定命令(
|
|
||||||
"method-work",
|
|
||||||
"work",
|
|
||||||
"work",
|
|
||||||
mid,
|
|
||||||
str(env["version"]["version_id"]),
|
|
||||||
),
|
|
||||||
)
|
|
||||||
with env["pools"][用途.生产].连接(只读=True) as conn:
|
|
||||||
material = 读取绑定材料(conn, "method-work", 内容用途="generation", 运行用途="production")
|
|
||||||
assert len(material) == 1
|
|
||||||
# 写手凭据不能扩为规划用途;拒绝码由原来的用途复核前移到凭据范围守卫。
|
|
||||||
with pytest.raises(Muse错误, match="不能扩为新目标或其他用途"):
|
|
||||||
读取绑定材料(conn, "method-work", 内容用途="planning", 运行用途="production")
|
|
||||||
usages = methods.列出消费(owner, str(env["version"]["version_id"]))
|
|
||||||
assert len(usages) == 1 and usages[0]["kind"] == "enablement"
|
|
||||||
assert usages[0]["result_ref"] == "B10.enablement:" + rid
|
|
||||||
before = len(env["received"])
|
|
||||||
svc.取消实验(env["actor"], certificate["result"]["experiment_id"], "stop-calibration")
|
|
||||||
with env["pools"][用途.生产].连接(只读=True) as conn:
|
|
||||||
assert 读取启用凭据(conn, owner.作者, rid)["stopped"]
|
|
||||||
with pytest.raises(Muse错误, match="停止"):
|
|
||||||
读取绑定材料(conn, "method-work", 内容用途="generation", 运行用途="production")
|
|
||||||
methods.变更方法状态(
|
|
||||||
owner, "disable-tested-method", 状态命令(mid, "disable", expected_state_revision=1)
|
|
||||||
)
|
|
||||||
with pytest.raises(Muse错误, match="停止"):
|
|
||||||
methods.变更方法状态(
|
|
||||||
owner,
|
|
||||||
"reenable-stopped-method",
|
|
||||||
状态命令(mid, "enable", expected_state_revision=1, 启用凭据=rid),
|
|
||||||
)
|
|
||||||
with pytest.raises(评测错误):
|
|
||||||
export()
|
|
||||||
assert methods.列出消费(owner, str(env["version"]["version_id"])) == usages
|
|
||||||
assert len(env["received"]) == before == 85
|
|
||||||
|
|
||||||
|
|
||||||
@pytest.mark.case_id(
|
@pytest.mark.case_id(
|
||||||
"NC-w25-25f902",
|
"NC-w25-25f902",
|
||||||
environment="隔离PG;85次合成HTTP及测试runtime验证器,不认证真实外部模型效果",
|
environment="隔离PG;无效凭据与角色权限,不发起合成评测循环",
|
||||||
given="确切方法确认版本和明确作者身份",
|
given="确切方法确认版本和明确作者身份",
|
||||||
when="核对启用凭据并执行当前目标的作者状态操作",
|
when="核对启用凭据并执行当前目标的作者状态操作",
|
||||||
then=["不存在或格式错误的凭据不能启用,生产和评测角色不能写凭据原表"],
|
then=["不存在或格式错误的凭据不能启用,生产和评测角色不能写凭据原表"],
|
||||||
@ -294,97 +100,3 @@ def test_未有目标证明不能启用且公开表不可写__25f902(方法环
|
|||||||
"SELECT has_table_privilege(current_user,%s,'SELECT')",
|
"SELECT has_table_privilege(current_user,%s,'SELECT')",
|
||||||
("public.muse_enablement_status",),
|
("public.muse_enablement_status",),
|
||||||
).fetchone()[0]
|
).fetchone()[0]
|
||||||
|
|
||||||
|
|
||||||
def _浏览器启用(env, config, rid, tmp_path):
|
|
||||||
import socket
|
|
||||||
import subprocess
|
|
||||||
from contextlib import asynccontextmanager
|
|
||||||
from pathlib import Path
|
|
||||||
from threading import Event, Thread
|
|
||||||
|
|
||||||
import uvicorn
|
|
||||||
|
|
||||||
from muse.接入.http.应用 import 创建应用
|
|
||||||
from muse.配置 import 服务配置, 读取配置
|
|
||||||
|
|
||||||
password = tmp_path / "method-browser-key"
|
|
||||||
password.write_text("synthetic-evaluation-only")
|
|
||||||
password.chmod(0o600)
|
|
||||||
sock = socket.socket()
|
|
||||||
sock.bind(("127.0.0.1", 0))
|
|
||||||
port = sock.getsockname()[1]
|
|
||||||
base = f"http://127.0.0.1:{port}"
|
|
||||||
app = 创建应用(
|
|
||||||
replace(
|
|
||||||
读取配置(config),
|
|
||||||
HTTP=服务配置(
|
|
||||||
str(password),
|
|
||||||
作者ID=env["owner"].作者,
|
|
||||||
端口=port,
|
|
||||||
公开地址=base,
|
|
||||||
允许来源=(base,),
|
|
||||||
),
|
|
||||||
)
|
|
||||||
)
|
|
||||||
previous = app.router.lifespan_context
|
|
||||||
ready = Event()
|
|
||||||
|
|
||||||
@asynccontextmanager
|
|
||||||
async def lifespan(instance):
|
|
||||||
async with previous(instance):
|
|
||||||
ready.set()
|
|
||||||
yield
|
|
||||||
|
|
||||||
app.router.lifespan_context = lifespan
|
|
||||||
server = uvicorn.Server(
|
|
||||||
uvicorn.Config(
|
|
||||||
app,
|
|
||||||
log_level="warning",
|
|
||||||
access_log=False,
|
|
||||||
timeout_graceful_shutdown=3,
|
|
||||||
)
|
|
||||||
)
|
|
||||||
thread = Thread(target=server.run, kwargs={"sockets": [sock]}, daemon=True)
|
|
||||||
output = Path(os.environ.get("MUSE_BROWSER_ARTIFACTS", str(tmp_path / "browser"))).resolve()
|
|
||||||
output.mkdir(parents=True, exist_ok=True)
|
|
||||||
process_env = {
|
|
||||||
**os.environ,
|
|
||||||
"MUSE_WORKBENCH_URL": base,
|
|
||||||
"MUSE_AUTHOR_PASSWORD_FILE": str(password),
|
|
||||||
"MUSE_METHOD_ID": env["method_id"],
|
|
||||||
"MUSE_ENABLEMENT_RECEIPT": rid,
|
|
||||||
}
|
|
||||||
chrome = Path("/Applications/Google Chrome.app/Contents/MacOS/Google Chrome")
|
|
||||||
if chrome.exists():
|
|
||||||
process_env.setdefault("MUSE_BROWSER_EXECUTABLE", str(chrome))
|
|
||||||
thread.start()
|
|
||||||
try:
|
|
||||||
assert ready.wait(10)
|
|
||||||
with (output / "浏览器.json").open("w") as log:
|
|
||||||
result = subprocess.run(
|
|
||||||
[
|
|
||||||
"pnpm",
|
|
||||||
"exec",
|
|
||||||
"playwright",
|
|
||||||
"test",
|
|
||||||
"方法启用.spec.ts",
|
|
||||||
"--reporter=json",
|
|
||||||
"--output",
|
|
||||||
str(output / "产物"),
|
|
||||||
],
|
|
||||||
cwd=Path(__file__).parents[2] / "web",
|
|
||||||
env=process_env,
|
|
||||||
stdout=log,
|
|
||||||
stderr=subprocess.STDOUT,
|
|
||||||
timeout=90,
|
|
||||||
)
|
|
||||||
assert result.returncode == 0, str(output / "浏览器.json")
|
|
||||||
receipts = list(output.rglob("启用回执.json"))
|
|
||||||
assert len(receipts) == 1
|
|
||||||
return json.loads(receipts[0].read_text())
|
|
||||||
finally:
|
|
||||||
server.should_exit = True
|
|
||||||
thread.join(timeout=10)
|
|
||||||
sock.close()
|
|
||||||
assert not thread.is_alive()
|
|
||||||
|
|||||||
@ -4,7 +4,6 @@ from decimal import Decimal
|
|||||||
|
|
||||||
import pytest
|
import pytest
|
||||||
import test_ABC正文回放 as ABC测试
|
import test_ABC正文回放 as ABC测试
|
||||||
import test_实际效果判据 as 效果测试
|
|
||||||
import test_文学评分执行与条件第三 as 文学测试
|
import test_文学评分执行与条件第三 as 文学测试
|
||||||
import test_标定凭据消费 as 消费测试
|
import test_标定凭据消费 as 消费测试
|
||||||
import test_标定金标准与凭据 as 标定测试
|
import test_标定金标准与凭据 as 标定测试
|
||||||
@ -126,7 +125,6 @@ def test_独立评委分歧逐维保留且不生成启用许可__254003(执行
|
|||||||
"真实浏览器从同一报告读回,刷新不追加模型调用",
|
"真实浏览器从同一报告读回,刷新不追加模型调用",
|
||||||
"资格报告显示实际标定覆盖和来源链接,读取不新增模型调用或启用",
|
"资格报告显示实际标定覆盖和来源链接,读取不新增模型调用或启用",
|
||||||
"实际检测发现和原字引文在报告页读回,未知与未配置分别显示",
|
"实际检测发现和原字引文在报告页读回,未知与未配置分别显示",
|
||||||
"实际效果分层、判据理由与凭据经工作台读回和封存;正式启用仍未评定",
|
|
||||||
],
|
],
|
||||||
contract="docs/系统架构/新版设计/模块设计/B10-效果评测.md",
|
contract="docs/系统架构/新版设计/模块设计/B10-效果评测.md",
|
||||||
)
|
)
|
||||||
@ -144,7 +142,6 @@ def test_独立评委分歧逐维保留且不生成启用许可__254003(执行
|
|||||||
("calibration", 标定测试.参数),
|
("calibration", 标定测试.参数),
|
||||||
("qualification", 标定测试.参数),
|
("qualification", 标定测试.参数),
|
||||||
("detection", 检测测试.参数),
|
("detection", 检测测试.参数),
|
||||||
("effect", 效果测试.参数),
|
|
||||||
],
|
],
|
||||||
indirect=["执行环境"],
|
indirect=["执行环境"],
|
||||||
)
|
)
|
||||||
@ -168,9 +165,6 @@ def test_浏览器报告来自同一隔离实验且读取不追加调用__254004
|
|||||||
if os.environ.get("MUSE_RUN_BROWSER") != "1":
|
if os.environ.get("MUSE_RUN_BROWSER") != "1":
|
||||||
pytest.skip("浏览器未显式启用,不计为通过")
|
pytest.skip("浏览器未显式启用,不计为通过")
|
||||||
env = request.getfixturevalue("ABC环境") if report_kind == "abc" else 执行环境
|
env = request.getfixturevalue("ABC环境") if report_kind == "abc" else 执行环境
|
||||||
if report_kind == "effect":
|
|
||||||
cert = 效果测试._校准(env)
|
|
||||||
env = 效果测试._新实验(env, 10, certificate=cert)
|
|
||||||
if report_kind == "qualification":
|
if report_kind == "qualification":
|
||||||
cert = 消费测试._校准(env)
|
cert = 消费测试._校准(env)
|
||||||
env = 消费测试._新实验(env, 消费测试._请求(env, [cert]))
|
env = 消费测试._新实验(env, 消费测试._请求(env, [cert]))
|
||||||
@ -181,7 +175,6 @@ def test_浏览器报告来自同一隔离实验且读取不追加调用__254004
|
|||||||
"calibration",
|
"calibration",
|
||||||
"qualification",
|
"qualification",
|
||||||
"detection",
|
"detection",
|
||||||
"effect",
|
|
||||||
}:
|
}:
|
||||||
env["scripted"].extend(["ok", "unknown_cost"])
|
env["scripted"].extend(["ok", "unknown_cost"])
|
||||||
运行测试._启动执行(env)
|
运行测试._启动执行(env)
|
||||||
@ -195,7 +188,6 @@ def test_浏览器报告来自同一隔离实验且读取不追加调用__254004
|
|||||||
"calibration",
|
"calibration",
|
||||||
"qualification",
|
"qualification",
|
||||||
"detection",
|
"detection",
|
||||||
"effect",
|
|
||||||
}:
|
}:
|
||||||
env["app"].要求评测().推进实验(env["actor"], env["exp"]["experiment_id"])
|
env["app"].要求评测().推进实验(env["actor"], env["exp"]["experiment_id"])
|
||||||
if report_kind == "corrected":
|
if report_kind == "corrected":
|
||||||
@ -205,7 +197,7 @@ def test_浏览器报告来自同一隔离实验且读取不追加调用__254004
|
|||||||
if report_kind == "detection":
|
if report_kind == "detection":
|
||||||
env["scripted"].extend(["high", "ok"])
|
env["scripted"].extend(["high", "ok"])
|
||||||
运行测试._运行就绪(env)
|
运行测试._运行就绪(env)
|
||||||
if report_kind in {"literary", "calibration", "qualification", "detection", "effect"}:
|
if report_kind in {"literary", "calibration", "qualification", "detection"}:
|
||||||
env["app"].要求评测().推进实验(env["actor"], env["exp"]["experiment_id"])
|
env["app"].要求评测().推进实验(env["actor"], env["exp"]["experiment_id"])
|
||||||
运行测试._运行就绪(env)
|
运行测试._运行就绪(env)
|
||||||
before = _报告(env)
|
before = _报告(env)
|
||||||
@ -260,7 +252,6 @@ def test_浏览器报告来自同一隔离实验且读取不追加调用__254004
|
|||||||
"calibration",
|
"calibration",
|
||||||
"qualification",
|
"qualification",
|
||||||
"detection",
|
"detection",
|
||||||
"effect",
|
|
||||||
}
|
}
|
||||||
else "0",
|
else "0",
|
||||||
}
|
}
|
||||||
@ -299,10 +290,6 @@ def test_浏览器报告来自同一隔离实验且读取不追加调用__254004
|
|||||||
assert before["calibration"]["receipt_id"] is None
|
assert before["calibration"]["receipt_id"] is None
|
||||||
assert after["calibration"]["receipt_id"]
|
assert after["calibration"]["receipt_id"]
|
||||||
assert before["samples"] == after["samples"]
|
assert before["samples"] == after["samples"]
|
||||||
elif report_kind == "effect":
|
|
||||||
assert before["effect"]["receipt_id"] is None and after["effect"]["receipt_id"]
|
|
||||||
assert before["effect"]["assessment"] == after["effect"]["assessment"]
|
|
||||||
assert before["samples"] == after["samples"]
|
|
||||||
else:
|
else:
|
||||||
assert after == before
|
assert after == before
|
||||||
assert len(env["received"]) == calls_before
|
assert len(env["received"]) == calls_before
|
||||||
|
|||||||
Loading…
x
Reference in New Issue
Block a user