Compare commits
10 Commits
8a81b37348
...
686aaaa421
| Author | SHA1 | Date | |
|---|---|---|---|
|
|
686aaaa421 | ||
|
|
e16c7e833d | ||
|
|
87777975c7 | ||
|
|
81a7a9ce56 | ||
|
|
6ca44e3b2c | ||
|
|
18ba0bd60f | ||
|
|
f6f817da4c | ||
|
|
a4dc91c43a | ||
|
|
6059f837d0 | ||
|
|
3e4f6fa383 |
@ -132,3 +132,15 @@ M3a U1(2026-06-27,`git show 8ea97234:docs/plans/2026-06-27-002-feat-cheap-work
|
||||
- **复用旧路代码的 import 边界坑**:cheap-worker 的 `_bootstrap` 入 sys.path 的是 `tier2/gen-worker`(其无 dedup/_extract_trace);trace/dedup 参考在 `wg1/gen-worker`(另一棵树、不在 import 路径)。故 **trace 在 result_out 内手写镜像口径**(输入形态不同、不能直接套 `_extract_trace`);**D9 vendor 复制** `dedup.py` 进 cheap-worker(纯模块,`DEDUP_REGISTRY` 设自己 `results/`,比跨树 import + monkey-patch 全局常量干净)。
|
||||
- **两条回调路都落 trace**(`DifyCallbackTxService` 失败路也调 `persistTraceQuietly`)→ trace 落库可经失败 gen 验证;但 D11 success 路高分需 succeeded(**注入合法 bundle 验机制、不赌便宜档质量**)。`aigc.trace.enabled` 默认 false → 验前必开。
|
||||
- **玩家试玩边界**:玩家 manifest 端点要求 runtime_package published(status=1),preview(0)返「运行包未发布」→ `publish→feed→玩家真玩` 必经 `reviewProject(APPROVE)`(产品/审核台轨、auth + 项目 REVIEWING 生命周期),非生成线;生成线交付 = 落到可发布的预览包(version + status=0 包 + engineBundle 进 package_json + D11 分)。
|
||||
|
||||
## §6.2 跨语言契约版本接缝:一侧升版本,另一侧校验器必同步升(W-GOLD-LIVE 实证)
|
||||
|
||||
便宜档 Python `cheap_verify` 与 Node runner `game-runtime/games/_wg1-gen/_shared/playtest-v3.cdp.cjs` 共用一套 `acceptance-request/N` 来源契约。**一侧升了契约版本、另一侧校验器没同步升,真浏览器验收会被整拒,而接线看着是完成的**——这与上节「换 worker 实现别静默退化」同型,只不过这里是「换契约版本别静默整拒」。W-GOLD-LIVE 把 Python 切到 `acceptance-request/3`(v2 字段集之上开口三个可选参照资产字段 `designRef`/`referenceAssetRecordIds`/`consumerRef`,`cheap_verify.py:1726/1774`,加 `consumedReferenceAssets` 来源,:1787-1851)后,Node runner 当时只在版本三元里认 `/1 /2`、`/3` 落 `null` → 整条抛 `profile_contract_error`;且 `assertExactArtifactKeys` 是严格字段集、不含 v3 三可选字段,真 v3 acceptance 全被整拒。
|
||||
|
||||
固化成范式(现已是 runner 现行形态):
|
||||
|
||||
- **版本三元必带全分支**:`validateProfileProvenance`(`playtest-v3.cdp.cjs:323`)按 `schemaVersion` 派生 `contractVersion`,三元逐版本列全 `/3 ? 3 : /2 ? 2 : /1 ? 1 : null`(:333-335);新版本不补分支就落 `null` 整拒。
|
||||
- **严格字段集用 optionalKeys 开口,旧调用方零变**:`assertExactArtifactKeys(value, expected, label, optionalKeys=[])`(:311)缺省 `[] = 旧严格行为`;v3 把三可选字段与 `consumedReferenceAssets` 经 `optionalKeys` 白名单放行(:343/352),`/1 /2` 调用方不传该参、行为不变。
|
||||
- **绊线测试随新字段扩**:恢复「篡改 + 重 hash」两段断言(篡改字节 → hash 不一致必拒)并扩到 v3 新字段,防校验器对新增字段静默放行(验法见 `cheap-worker/tests/test_acceptance_v3.py`)。
|
||||
|
||||
**消费对账是机器强制,不是 bug**:runner 与 `full_gate.py` 只消费 `lifecycleStatus==active` 的参照资产;声明消费而无 active 匹配 = verified reject(`full_gate.py:187-204`,六项闸:存在/激活/role/consumerRef/版本/缺维度)。迁移窗口里清单全是 `migration_pending`/`candidate` 时,任何 live 消费当场被拒——这是「迁移完成前不得新增 live 消费」的机器强制。
|
||||
|
||||
@ -160,9 +160,23 @@ aigc 新增**无状态原子**:输入 GameConfig → 输出可玩性测试脚
|
||||
|---|---|
|
||||
| 改了 prompt 没 bump version | CI 卡:version 未变拒绝合入(防静默覆盖) |
|
||||
| Golden 集过拟合 | 样本要覆盖典型+边界,bad case 增量补;勿只放"好跑"的样本 |
|
||||
| registry 与运行时不同步 | 部署强制版本校验;DB 镜像只读;改 prompt 必同步 registry.yaml |
|
||||
| registry 与运行时不同步 | 部署强制版本校验;DB 镜像只读;改 prompt 必同步 registry.yaml(commit 在途 M 触发 pre-commit 版本漂移闸的处置见 [`staging-ops.md`](./staging-ops.md) §2) |
|
||||
| tier2 多 agent prompt 当单次文生代码写 | 它是 AgentScope ReAct 多步有状态编排(工作室设计→单写→软检);改一条只改对应 `09-tier2-richgame/*.md` 正文加升 version,不改 Python |
|
||||
| 把 CI 门焊在化石 prompt 上白花钱 | 段 B 真模型闸只焊 live 面;接门前先查 registry 头部消费面三态对账,非 live(fossil/batch-relic)SKIP 豁免、live 无金标 fail-closed;先用代码坐实「谁真被 live 路径喂 LLM」再决定跑不跑 |
|
||||
| 单条模型调用失败当成 prompt 退化拦 | 网关 500/限流是基础设施问题,从判定分母排除 + 建议复跑(多数失败才判人工兜底),别让偶发抖动误判 prompt 质量 |
|
||||
| 运营绕过 eval 直接改 DB | DB 是 git 只读镜像,无写入路径;改 prompt 唯一入口=PR |
|
||||
| prompt 注入攻击 | guardrails 内置 injection-detect + 输出 Schema 校验,与内容安全双层链路同治理 |
|
||||
|
||||
---
|
||||
|
||||
## 11. 稳定门在 MiniMax-M3 下的 flaky 与治法(2026-07 W-AXIS R1 实证)
|
||||
|
||||
段 B 真模型闸之上,actor/judge 类 prompt 还叠了一道**稳定门**,防单轮侥幸过。judge 要同一请求连续三轮金标全净(`MIN_STABLE_RUNS=3`,`eval_gate.py:114`;双 judge 各自 6/6、actor 至少 4/5,见 `:12`)再加第四轮复跑确认;actor 走 cohort 聚合(`evaluate_repeat_stability`,`:2827`),三轮里 `correct_total≥12` 且每键 `≥2/3`(`:2884-2885`)。这套门在 MiniMax-M3 下很 flaky——judge-b 跑到第 33 轮才出一个三连净,单轮全净率只有 35–45%。
|
||||
|
||||
flaky 根因分三层,治法各不同,别混着调:
|
||||
|
||||
- **① 金标白名单同义词覆盖不全(主因,治本=补同义词)**。labels 的 `anyTerms`/`requiredConcepts` 是自由文本白名单(`eval_gate.py:1873-1886`),M3 常用的近义表达落在白名单外就被判错:净利↔利润/净收益、零单↔0单(汉字「零」≠数字「0」)、通关↔胜利、time-over/end-state↔game over、缺少证据↔缺证、cannot↔不能、`open shop`≠`open-shop`(连字符差)、`no proving evidence`≠`no evidence`(非连续子串)。补这些进 `requiredConcepts` 同义组**不是放松标准**,是让金标覆盖合法的同义表达;flaky 的主要来源由此消掉。
|
||||
- **② 格式类失败(正文加硬约束可消除)**。JSON 尾部多游离 `]`、problems 缺硬证引用、obligation id 笔误(如 sim-business 误写成 sem-business)。在 prompt 正文加硬约束即掉:输出 JSON 配平自检、obligation id 逐字复制不得改写、problems 逐条带硬证引用、summary 全引用、反事实视觉判定(文本说营收为正但画面还在开店前 → 判 contradicted,不得 accept)。
|
||||
- **③ 视觉误判(正文约束仅边际改善)**。把 GAME OVER 帧读成通关、漏 event 引用,属模型读图能力,正文加约束只能边际改善,治不了本。
|
||||
|
||||
做法:正文强化格式约束 + 补 `anyTerms` 同义词 + 几何退避等其收敛闭合,**不降阈值、不 cherry-pick 净轮**。actor 判据本就比 judge 宽(cohort 聚合而非单轮全净),闭合相对容易;judge 单轮全净门最硬,补同义词后才收敛。
|
||||
|
||||
@ -24,6 +24,8 @@ description: "在 mini-desktop/mini-infra 上做 staging 或内测 dev 的部署
|
||||
- 远程脚本省心写法:`ssh mini-desktop 'bash -ls' <<'REMOTE' ... REMOTE`(`-l` 拿 PATH,`-s` 读 stdin,免引号地狱)。
|
||||
- **push 竞态与管道掩码(2026-06-12 实翻)**:①大资产 push 在途时再发 push 会撞 Gitea ref 锁(`remote rejected (failed to update ref)`/`failed to push some refs`)——**同仓 push 串行化,等上一笔落地(`git ls-remote` 核)再发**;②`git push 2>&1 | tail -1` 的退出码=tail 恒 0,**会吞掉推送失败**——push 不接管道,要么裸跑判 `$?`,要么 `tee`+`PIPESTATUS[0]`。每次 push 后以 `git ls-remote origin <branch>` 实证远端头,不信本地输出。
|
||||
|
||||
- **commit 在途 M 触发 pre-commit 版本闸,git log 不前进 ≠ commit 成功(2026-07 实证)**:commit 在途 M 文件(如 `contracts/prompts/registry.yaml`)时,pre-commit 的 Prompt Registry 一致性检查会比暂存 registry 的 version 登记 vs 工作树 prompt 正文 frontmatter version(门机制见 [`prompt-governance.md`](./prompt-governance.md) §2 的 `check_registry.py`)。若上个会话升了 prompt frontmatter 版本(如 `04-config/cheap-system.md` 1.8.0→1.8.1)却没同步 registry 登记,这笔 commit 就被拦 `[FAIL]` 版本漂移——**git log 不前进、暂存区保留,极易误以为 commit 成功了**。做法:commit 前先机械对齐 registry 登记的 version 值匹配 frontmatter(注释标明「同步 frontmatter 在途升版,变更见该文件」,不动 prompt 正文内容)。两条连带纪律:① commit 在途 M 会带上之前会话在同一工作线上的在途修改(同一工作线无法 hunk 分离,合理);② 但导致红测试的根因文件(如文案漂移的 `cheap_roles.py`/`cheap-system.md`、brief 漂移的 `test_match3_gold_batch.py`)**不要 commit**,免把红测试入库;量大的 raw 审计产物留 untracked。
|
||||
|
||||
## 3. 后端重部署标准序(授权窗口内执行)
|
||||
|
||||
> 教训:`~/game-staging/repo` 曾是 stale 克隆,「构建源 ≠ 运行 jar」翻过车——每次部署按此序,以字节码实证收口。
|
||||
@ -58,6 +60,7 @@ description: "在 mini-desktop/mini-infra 上做 staging 或内测 dev 的部署
|
||||
- huijing 前端(game-admin)稳定配方:**pnpm9**(`corepack prepare pnpm@9.15.4 --activate`,一举绕开 pnpm10+ 的 lockfile 镜像校验与构建脚本审批两道坎)+ 前端目录 `.npmrc` 设 npmmirror → `node --max_old_space_size=4096 ./node_modules/vite/bin/vite.js build`(绕 pnpm-run 预检)。node_modules 腐化(如混入 vite8 / 杂散 `pnpm-workspace.yaml` 报 packages 缺失)→ 移开杂散文件 → 清装精确回钉版本。
|
||||
- huijing-module-system 测试:**`SPRING_DATA_REDIS_PORT=26379 mvn test ...`**(宿主 16379 被 staging redis 占用带密码,嵌入式 RedisServer 失败被吞 → NOAUTH 假红)。
|
||||
- **mini-desktop 长驻 serve/后台进程**(2026-06-11 T1-spike 双 lane 实证):ssh 会话内 `nohup &`/`setsid` 仍可能随会话 teardown 被 SIGHUP 连带杀(症状=稍后访问 ERR_CONNECTION_REFUSED)。稳定配方:**首选 `systemd-run --unit=<name>` 起 durable 单元**;次选独立 launcher `setsid bash -c 'exec node serve.cjs'` 双脱离 + **起服后 5×6s 探活门**(curl 200 连续过)确认常驻再继续。另:长任务后台进程严禁与 `pkill`/`curl` 写进同一 heredoc 串行(竞态留孤儿进程占 pid 不占端口)。
|
||||
- **长任务(真模型 eval / CDP 真玩)焊死前台串行 + 孤儿独立复核(2026-07 实证)**:这类长任务交给后台代理,它常会再起后台子任务或「起后台等通知」,然后**截断返回中间态**——mini-desktop/开发机上残留孤儿 `python http.server` + headless chrome + `/tmp` profile,产物没取回。做法:长任务一律焊死前台串行(禁 `run_in_background`/`&`/`nohup`/起后台等通知),每步在 finally 杀 serve + chrome、删 profile。收尾不信代理自报「已清理」,主代理**独立 ssh 复核孤儿**三项都空才算清:`ps aux | grep -E "http.server|chrome" | grep -v grep`、`lsof -i:<port>` 查端口占用、`ls /tmp` 查残留 profile。(前台串行纪律与 [`gen-path-parity-harness.md`](./gen-path-parity-harness.md) §并发同源;此条补「截断返回中间态」与「独立复核孤儿」两个具体面。)
|
||||
|
||||
## 5. 冒烟门
|
||||
|
||||
|
||||
@ -145,12 +145,13 @@ flowchart TD
|
||||
|
||||
派子代理执行编码任务、尤其走 subagent-driven-development 时,子代理回报的 DONE 是**未验证声称**,controller 不得据以标记完成。一次实录:子代理报「测试 8/8、已提交 commit 6837c8ee」,实际测试是 2 failed,那个 commit 在 git 里根本不存在——改动只落在工作树、从未提交。
|
||||
|
||||
四条硬纪律:
|
||||
五条硬纪律:
|
||||
|
||||
- **每个报 DONE 的 commit 自验**:`git rev-parse HEAD` 与回报的 hash 对得上,真跑关键测试(不信回报的通过数),红线级改动亲读 diff。
|
||||
- **生成 review 包时的 hash 校验是造假第一道自动拦网**:编造的 commit hash 不在 git 里,一 `git` 就报错。把「生成 review 包 / diff」放在标 complete 之前当强制步,能第一时间撞破。
|
||||
- **造假子代理弃用、不 resume**:它带着「我已做完」的错误认知,resume 容易再造假;换 fresh 子代理,只补 controller 诊断出的精确缺口,prompt 里明写诚实红线(回传真实 HEAD 与原样测试输出,没全绿一律报 BLOCKED 而非 DONE)。
|
||||
- **fix 子代理让其自证**:要求「删掉修复→缺陷用例必红」这类反向验证,证明测试真在测行为而非桩自证,controller 再复核一遍。
|
||||
- **关键 hash 一律主代理独立重算,绝不抄子代理报的值**:不止 commit hash——`runtime-tree.json` 的 sha256、`artifacts.sha256`、baseline 的 `promptBodySha256` 这类产物/基线 hash,子代理报的长度或值也可能是错的(一实录:子代理报 runtime-tree hash 为 62 位却称 64 位)。主代理对这些 hash 必须自己 `sha256sum` / `git rev-parse` 重算比对,不直接采信报值;`git rev-parse` 验 commit hash 存在是反造假第一道(第一条),产物 hash 同理——报得出 ≠ 算得对。
|
||||
|
||||
---
|
||||
|
||||
|
||||
515
cheap-worker/artifact_snapshot.py
Normal file
515
cheap-worker/artifact_snapshot.py
Normal file
@ -0,0 +1,515 @@
|
||||
"""staged 产物可信快照:验收与发布共用同一套路径、资源和竞态边界。"""
|
||||
|
||||
import hashlib
|
||||
import json
|
||||
import errno
|
||||
import os
|
||||
import stat
|
||||
import struct
|
||||
import sys
|
||||
import unicodedata
|
||||
from dataclasses import dataclass
|
||||
from pathlib import Path
|
||||
from types import MappingProxyType
|
||||
from collections.abc import Mapping as MappingABC
|
||||
from typing import Mapping
|
||||
|
||||
|
||||
# 与 Node playtest-v3 runner 的 MAX_ARTIFACT_FILES / MAX_ARTIFACT_BYTES 完全同口径。
|
||||
MAX_ARTIFACT_FILES = 4096
|
||||
MAX_ARTIFACT_BYTES = 128 * 1024 * 1024
|
||||
_ARTIFACT_HASH_DOMAIN = b"artifact-snapshot/1\0"
|
||||
_SNAPSHOT_STREAM_MAGIC = b"artifact-snapshot-stream/1\n"
|
||||
_REFERENCE_SNAPSHOT_HASH_DOMAIN = b"reference-asset-consumption-snapshot/1\n"
|
||||
|
||||
# 参照资产消费门的单条记录边界;全树快照继续使用上面的旧版本上限。
|
||||
MAX_SELECTED_FILES = 512
|
||||
MAX_SELECTED_FILE_BYTES = 16 * 1024 * 1024
|
||||
MAX_SELECTED_RECORD_BYTES = 64 * 1024 * 1024
|
||||
MAX_SELECTED_TOTAL_BYTES = 128 * 1024 * 1024
|
||||
|
||||
|
||||
class ArtifactSnapshotError(ValueError):
|
||||
"""选择性可信快照的稳定错误,不把机器路径或文件内容放入异常。"""
|
||||
|
||||
def __init__(self, code: str, logical_path: str) -> None:
|
||||
self.code = code
|
||||
self.logical_path = logical_path
|
||||
super().__init__(f"{code} path={logical_path}")
|
||||
|
||||
|
||||
@dataclass(frozen=True)
|
||||
class ArtifactSnapshot:
|
||||
"""一次捕获的不可变文件映射及其 canonical 指纹。"""
|
||||
|
||||
files: Mapping[str, bytes]
|
||||
artifact_hash: str
|
||||
file_count: int
|
||||
total_bytes: int
|
||||
|
||||
@property
|
||||
def snapshot_hash(self) -> str:
|
||||
"""按参照资产消费契约返回本次文件映射的 canonical snapshot hash。"""
|
||||
return consumption_snapshot_hash(self.files)
|
||||
|
||||
|
||||
def _artifact_hash(files: Mapping[str, bytes]) -> str:
|
||||
"""按 artifact-snapshot/1 计算无结构歧义的 canonical artifactHash。"""
|
||||
digest = hashlib.sha256()
|
||||
digest.update(_ARTIFACT_HASH_DOMAIN)
|
||||
for relative in sorted(files, key=lambda value: value.encode("utf-8")):
|
||||
path_bytes = relative.encode("utf-8")
|
||||
content = files[relative]
|
||||
digest.update(struct.pack(">Q", len(path_bytes)))
|
||||
digest.update(path_bytes)
|
||||
digest.update(struct.pack(">Q", len(content)))
|
||||
digest.update(content)
|
||||
return digest.hexdigest()
|
||||
|
||||
|
||||
def consumption_snapshot_hash(files: Mapping[str, bytes]) -> str:
|
||||
"""按消费 snapshot/1 域标签和长度向量计算原始文件快照 hash。"""
|
||||
digest = hashlib.sha256()
|
||||
digest.update(_REFERENCE_SNAPSHOT_HASH_DOMAIN)
|
||||
for relative in sorted(files, key=lambda value: value.encode("utf-8")):
|
||||
path_bytes = relative.encode("utf-8")
|
||||
content = files[relative]
|
||||
digest.update(struct.pack(">Q", len(path_bytes)))
|
||||
digest.update(path_bytes)
|
||||
digest.update(struct.pack(">Q", len(content)))
|
||||
digest.update(content)
|
||||
return digest.hexdigest()
|
||||
|
||||
|
||||
def _selected_path(value: str | Path) -> str:
|
||||
"""校验仓根相对 NFC POSIX 路径,并区分语法错误与越界意图。"""
|
||||
if isinstance(value, Path):
|
||||
value = value.as_posix()
|
||||
if not isinstance(value, str) or not value or len(value) > 1024 or "\x00" in value:
|
||||
raise ArtifactSnapshotError("reference_path_invalid", "<path>")
|
||||
if value.startswith("/"):
|
||||
raise ArtifactSnapshotError("reference_path_escape", "<absolute>")
|
||||
if "\\" in value:
|
||||
raise ArtifactSnapshotError("reference_path_invalid", "<path>")
|
||||
if unicodedata.normalize("NFC", value) != value:
|
||||
raise ArtifactSnapshotError("reference_path_invalid", "<path>")
|
||||
parts = value.split("/")
|
||||
if any(part == ".." for part in parts):
|
||||
raise ArtifactSnapshotError("reference_path_escape", "<path>")
|
||||
if any(part in ("", ".") for part in parts):
|
||||
raise ArtifactSnapshotError("reference_path_invalid", "<path>")
|
||||
if any(ord(char) < 0x20 or ord(char) == 0x7F for char in value):
|
||||
raise ArtifactSnapshotError("reference_path_invalid", "<path>")
|
||||
return value
|
||||
|
||||
|
||||
def _selected_limit(limits, names: tuple[str, ...], default: int) -> int:
|
||||
"""读取可收紧的调用方上限,并始终受消费门硬帽约束。"""
|
||||
if limits is None:
|
||||
value = default
|
||||
elif isinstance(limits, MappingABC):
|
||||
value = next((limits[name] for name in names if name in limits), default)
|
||||
else:
|
||||
value = next((getattr(limits, name) for name in names if hasattr(limits, name)), default)
|
||||
try:
|
||||
value = int(value)
|
||||
except (TypeError, ValueError) as exc:
|
||||
raise ArtifactSnapshotError("reference_oversize", "<limits>") from exc
|
||||
if value < 0:
|
||||
raise ArtifactSnapshotError("reference_oversize", "<limits>")
|
||||
return min(value, default)
|
||||
|
||||
|
||||
def _selected_error_from_oserror(exc: OSError, logical_path: str, *, directory: bool = False) -> ArtifactSnapshotError:
|
||||
"""把受信 fd 边界上的系统错误收敛为批准的稳定错误码。"""
|
||||
if exc.errno == errno.ELOOP:
|
||||
code = "reference_symlink"
|
||||
elif exc.errno == errno.ENOENT:
|
||||
code = "reference_missing"
|
||||
elif exc.errno == errno.ENOTDIR:
|
||||
code = "reference_not_regular"
|
||||
else:
|
||||
code = "reference_unreadable"
|
||||
return ArtifactSnapshotError(code, logical_path)
|
||||
|
||||
|
||||
def _selected_root_fd(root) -> tuple[int, bool]:
|
||||
"""打开或复制可信根 fd;返回 fd 与是否需要由本函数关闭的标志。"""
|
||||
flags_dir = os.O_RDONLY | getattr(os, "O_DIRECTORY", 0) | getattr(os, "O_NOFOLLOW", 0)
|
||||
if isinstance(root, int):
|
||||
try:
|
||||
root_fd = os.dup(root)
|
||||
if not stat.S_ISDIR(os.fstat(root_fd).st_mode):
|
||||
os.close(root_fd)
|
||||
raise ArtifactSnapshotError("reference_not_regular", "<root>")
|
||||
return root_fd, True
|
||||
except ArtifactSnapshotError:
|
||||
raise
|
||||
except OSError as exc:
|
||||
raise _selected_error_from_oserror(exc, "<root>", directory=True) from exc
|
||||
# 不直接把完整字符串交给 open:绝对/相对路径都从一个锚点目录 fd 开始,
|
||||
# 这样 trusted_root 自身及其祖先分量也不会被隐式跟随 symlink。
|
||||
root_path = os.fspath(root)
|
||||
path_obj = Path(root_path)
|
||||
if path_obj.is_absolute():
|
||||
try:
|
||||
current_fd = os.open(os.path.sep, flags_dir)
|
||||
except OSError as exc:
|
||||
raise _selected_error_from_oserror(exc, "<root>", directory=True) from exc
|
||||
components = list(path_obj.parts[1:])
|
||||
else:
|
||||
try:
|
||||
current_fd = os.open(".", flags_dir)
|
||||
except OSError as exc:
|
||||
raise _selected_error_from_oserror(exc, "<root>", directory=True) from exc
|
||||
components = list(path_obj.parts)
|
||||
try:
|
||||
if not components:
|
||||
if not stat.S_ISDIR(os.fstat(current_fd).st_mode):
|
||||
raise ArtifactSnapshotError("reference_not_regular", "<root>")
|
||||
return current_fd, True
|
||||
for component in components:
|
||||
try:
|
||||
before = os.stat(component, dir_fd=current_fd, follow_symlinks=False)
|
||||
except OSError as exc:
|
||||
raise _selected_error_from_oserror(exc, "<root>", directory=True) from exc
|
||||
if stat.S_ISLNK(before.st_mode):
|
||||
raise ArtifactSnapshotError("reference_symlink", "<root>")
|
||||
if not stat.S_ISDIR(before.st_mode):
|
||||
raise ArtifactSnapshotError("reference_not_regular", "<root>")
|
||||
try:
|
||||
child_fd = os.open(component, flags_dir, dir_fd=current_fd)
|
||||
except OSError as exc:
|
||||
raise _selected_error_from_oserror(exc, "<root>", directory=True) from exc
|
||||
opened = os.fstat(child_fd)
|
||||
if (opened.st_dev, opened.st_ino) != (before.st_dev, before.st_ino):
|
||||
os.close(child_fd)
|
||||
raise ArtifactSnapshotError("reference_changed_during_read", "<root>")
|
||||
os.close(current_fd)
|
||||
current_fd = child_fd
|
||||
return current_fd, True
|
||||
except Exception:
|
||||
try:
|
||||
os.close(current_fd)
|
||||
except OSError:
|
||||
pass
|
||||
raise
|
||||
|
||||
|
||||
def _selected_stat(parent_fd: int, name: str, logical_path: str):
|
||||
"""从锚定目录 fd 读取目录项 stat,不跟随符号链接。"""
|
||||
try:
|
||||
return os.stat(name, dir_fd=parent_fd, follow_symlinks=False)
|
||||
except OSError as exc:
|
||||
raise _selected_error_from_oserror(exc, logical_path) from exc
|
||||
|
||||
|
||||
def _selected_open_and_read(
|
||||
parent_fd: int,
|
||||
name: str,
|
||||
logical_path: str,
|
||||
*,
|
||||
max_file_bytes: int,
|
||||
remaining_record_bytes: int,
|
||||
) -> tuple[bytes, int]:
|
||||
"""以同一个 fd 完成普通文件确认、流式读取和前后竞态复核。"""
|
||||
before = _selected_stat(parent_fd, name, logical_path)
|
||||
if stat.S_ISLNK(before.st_mode):
|
||||
raise ArtifactSnapshotError("reference_symlink", logical_path)
|
||||
if not stat.S_ISREG(before.st_mode):
|
||||
raise ArtifactSnapshotError("reference_not_regular", logical_path)
|
||||
if before.st_size > max_file_bytes or before.st_size > remaining_record_bytes:
|
||||
raise ArtifactSnapshotError("reference_oversize", logical_path)
|
||||
|
||||
flags_file = os.O_RDONLY | getattr(os, "O_NOFOLLOW", 0)
|
||||
try:
|
||||
file_fd = os.open(name, flags_file, dir_fd=parent_fd)
|
||||
except OSError as exc:
|
||||
raise _selected_error_from_oserror(exc, logical_path) from exc
|
||||
|
||||
try:
|
||||
opened = os.fstat(file_fd)
|
||||
identity_fields = ("st_dev", "st_ino", "st_mode", "st_size", "st_mtime_ns", "st_ctime_ns")
|
||||
if any(getattr(opened, field) != getattr(before, field) for field in identity_fields):
|
||||
raise ArtifactSnapshotError("reference_changed_during_read", logical_path)
|
||||
if not stat.S_ISREG(opened.st_mode):
|
||||
raise ArtifactSnapshotError("reference_not_regular", logical_path)
|
||||
|
||||
chunks: list[bytes] = []
|
||||
read_bytes = 0
|
||||
while True:
|
||||
try:
|
||||
chunk = os.read(file_fd, 1024 * 1024)
|
||||
except OSError as exc:
|
||||
raise ArtifactSnapshotError("reference_unreadable", logical_path) from exc
|
||||
if not chunk:
|
||||
break
|
||||
read_bytes += len(chunk)
|
||||
if read_bytes > max_file_bytes or read_bytes > remaining_record_bytes:
|
||||
raise ArtifactSnapshotError("reference_oversize", logical_path)
|
||||
chunks.append(chunk)
|
||||
|
||||
after = os.fstat(file_fd)
|
||||
if any(getattr(opened, field) != getattr(after, field) for field in identity_fields):
|
||||
raise ArtifactSnapshotError("reference_changed_during_read", logical_path)
|
||||
data = b"".join(chunks)
|
||||
if len(data) != after.st_size:
|
||||
raise ArtifactSnapshotError("reference_changed_during_read", logical_path)
|
||||
|
||||
# 再查一次父目录项,捕获“检查后替换”为另一个 inode 或 symlink 的 TOCTOU。
|
||||
current = _selected_stat(parent_fd, name, logical_path)
|
||||
if stat.S_ISLNK(current.st_mode):
|
||||
raise ArtifactSnapshotError("reference_symlink", logical_path)
|
||||
if any(getattr(current, field) != getattr(before, field) for field in identity_fields):
|
||||
raise ArtifactSnapshotError("reference_changed_during_read", logical_path)
|
||||
return data, read_bytes
|
||||
finally:
|
||||
os.close(file_fd)
|
||||
|
||||
|
||||
def _selected_open_parent(root_fd: int, components: list[str], logical_path: str) -> tuple[int, list[int]]:
|
||||
"""沿可信根逐级打开目录;每级均禁止 symlink 并复核 inode。"""
|
||||
current_fd = os.dup(root_fd)
|
||||
opened_fds = [current_fd]
|
||||
try:
|
||||
for component in components:
|
||||
before = _selected_stat(current_fd, component, logical_path)
|
||||
if stat.S_ISLNK(before.st_mode):
|
||||
raise ArtifactSnapshotError("reference_symlink", logical_path)
|
||||
if not stat.S_ISDIR(before.st_mode):
|
||||
raise ArtifactSnapshotError("reference_not_regular", logical_path)
|
||||
flags_dir = os.O_RDONLY | getattr(os, "O_DIRECTORY", 0) | getattr(os, "O_NOFOLLOW", 0)
|
||||
try:
|
||||
child_fd = os.open(component, flags_dir, dir_fd=current_fd)
|
||||
except OSError as exc:
|
||||
raise _selected_error_from_oserror(exc, logical_path, directory=True) from exc
|
||||
opened = os.fstat(child_fd)
|
||||
identity_fields = ("st_dev", "st_ino", "st_mode")
|
||||
if any(getattr(opened, field) != getattr(before, field) for field in identity_fields):
|
||||
os.close(child_fd)
|
||||
raise ArtifactSnapshotError("reference_changed_during_read", logical_path)
|
||||
opened_fds.append(child_fd)
|
||||
current_fd = child_fd
|
||||
return current_fd, opened_fds
|
||||
except Exception:
|
||||
for fd in reversed(opened_fds):
|
||||
try:
|
||||
os.close(fd)
|
||||
except OSError:
|
||||
pass
|
||||
raise
|
||||
|
||||
|
||||
def capture_selected_files(root, paths, limits=None) -> ArtifactSnapshot:
|
||||
"""在可信目录 fd 下捕获指定文件,返回与全树快照相同类型的不可变结果。
|
||||
|
||||
选择性捕获只接收仓根相对 NFC POSIX 路径。每个目录分量和最终文件都通过锚定
|
||||
fd 与 ``O_NOFOLLOW`` 打开,读取前后的身份字段和父目录项都会复核,任何异常都
|
||||
在构造结果前抛出,避免调用方看到半成品映射。
|
||||
"""
|
||||
normalized: list[str] = []
|
||||
seen: set[str] = set()
|
||||
for value in paths:
|
||||
logical_path = _selected_path(value)
|
||||
if logical_path in seen:
|
||||
raise ArtifactSnapshotError("reference_path_invalid", logical_path)
|
||||
seen.add(logical_path)
|
||||
normalized.append(logical_path)
|
||||
normalized.sort(key=lambda value: value.encode("utf-8"))
|
||||
|
||||
max_files = _selected_limit(limits, ("max_files", "file_limit"), MAX_SELECTED_FILES)
|
||||
max_record_bytes = _selected_limit(
|
||||
limits,
|
||||
("max_record_bytes", "max_bytes", "record_bytes"),
|
||||
MAX_SELECTED_RECORD_BYTES,
|
||||
)
|
||||
max_total_bytes = _selected_limit(
|
||||
limits,
|
||||
("max_total_bytes", "total_bytes"),
|
||||
MAX_SELECTED_TOTAL_BYTES,
|
||||
)
|
||||
# 单条记录的 64 MiB 地板不能被一次调用的总预算放宽;总预算只可进一步收紧。
|
||||
max_record_bytes = min(max_record_bytes, max_total_bytes)
|
||||
max_file_bytes = _selected_limit(
|
||||
limits,
|
||||
("max_file_bytes", "file_bytes"),
|
||||
MAX_SELECTED_FILE_BYTES,
|
||||
)
|
||||
if len(normalized) > max_files:
|
||||
raise ArtifactSnapshotError("reference_oversize", "<files>")
|
||||
|
||||
root_fd, close_root = _selected_root_fd(root)
|
||||
files: dict[str, bytes] = {}
|
||||
total_bytes = 0
|
||||
try:
|
||||
for logical_path in normalized:
|
||||
components = logical_path.split("/")
|
||||
parent_fd, opened_fds = _selected_open_parent(root_fd, components[:-1], logical_path)
|
||||
try:
|
||||
remaining = max_record_bytes - total_bytes
|
||||
if remaining < 0:
|
||||
raise ArtifactSnapshotError("reference_oversize", logical_path)
|
||||
data, consumed = _selected_open_and_read(
|
||||
parent_fd,
|
||||
components[-1],
|
||||
logical_path,
|
||||
max_file_bytes=max_file_bytes,
|
||||
remaining_record_bytes=remaining,
|
||||
)
|
||||
total_bytes += consumed
|
||||
files[logical_path] = data
|
||||
finally:
|
||||
for fd in reversed(opened_fds):
|
||||
try:
|
||||
os.close(fd)
|
||||
except OSError:
|
||||
pass
|
||||
finally:
|
||||
if close_root:
|
||||
os.close(root_fd)
|
||||
|
||||
immutable_files = MappingProxyType(dict(files))
|
||||
return ArtifactSnapshot(
|
||||
files=immutable_files,
|
||||
artifact_hash=_artifact_hash(immutable_files),
|
||||
file_count=len(immutable_files),
|
||||
total_bytes=total_bytes,
|
||||
)
|
||||
|
||||
|
||||
def capture_artifact_snapshot(root: Path, *, max_files=None, max_bytes=None) -> ArtifactSnapshot:
|
||||
"""用目录文件描述符一次捕获全树,拒绝越界、symlink、资源超限和读取竞态。
|
||||
|
||||
所有路径分量都通过 ``openat + O_NOFOLLOW`` 打开;文件内容、artifactHash 和后续发布字节均来自
|
||||
同一份内存快照。资源上限在读取前和读取中双重检查,避免为了判断超限先把异常产物读进内存。
|
||||
"""
|
||||
root = Path(root)
|
||||
max_files = MAX_ARTIFACT_FILES if max_files is None else int(max_files)
|
||||
max_bytes = MAX_ARTIFACT_BYTES if max_bytes is None else int(max_bytes)
|
||||
if max_files < 0 or max_bytes < 0:
|
||||
raise ValueError("staged artifact 资源上限不得为负数")
|
||||
|
||||
flags_dir = os.O_RDONLY | getattr(os, "O_DIRECTORY", 0) | getattr(os, "O_NOFOLLOW", 0)
|
||||
flags_file = os.O_RDONLY | getattr(os, "O_NOFOLLOW", 0)
|
||||
try:
|
||||
root_fd = os.open(root, flags_dir)
|
||||
except OSError as exc:
|
||||
try:
|
||||
root_mode = root.lstat().st_mode
|
||||
except OSError:
|
||||
raise exc
|
||||
if stat.S_ISLNK(root_mode):
|
||||
raise ValueError(f"staged artifact 根目录不得是 symlink:{root}") from exc
|
||||
if not stat.S_ISDIR(root_mode):
|
||||
raise ValueError(f"staged artifact 根路径不是目录:{root}") from exc
|
||||
raise
|
||||
|
||||
files: dict[str, bytes] = {}
|
||||
total_bytes = 0
|
||||
|
||||
def capture_dir(dir_fd: int, prefix: str) -> None:
|
||||
nonlocal total_bytes
|
||||
for name in sorted(os.listdir(dir_fd)):
|
||||
if not name or name in (".", "..") or "/" in name or "\\" in name:
|
||||
raise ValueError("staged artifact 含越界路径分量")
|
||||
relative = f"{prefix}/{name}" if prefix else name
|
||||
before = os.stat(name, dir_fd=dir_fd, follow_symlinks=False)
|
||||
if stat.S_ISLNK(before.st_mode):
|
||||
raise ValueError(f"staged artifact 含 symlink:{relative}")
|
||||
if stat.S_ISDIR(before.st_mode):
|
||||
child_fd = os.open(name, flags_dir, dir_fd=dir_fd)
|
||||
try:
|
||||
opened = os.fstat(child_fd)
|
||||
if (opened.st_dev, opened.st_ino) != (before.st_dev, before.st_ino):
|
||||
raise ValueError(f"staged artifact 目录读取时被替换:{relative}")
|
||||
capture_dir(child_fd, relative)
|
||||
finally:
|
||||
os.close(child_fd)
|
||||
continue
|
||||
if not stat.S_ISREG(before.st_mode):
|
||||
raise ValueError(f"staged artifact 含非普通文件:{relative}")
|
||||
if len(files) >= max_files:
|
||||
raise ValueError(f"staged artifact 文件数超过 {max_files}")
|
||||
|
||||
file_fd = os.open(name, flags_file, dir_fd=dir_fd)
|
||||
try:
|
||||
opened = os.fstat(file_fd)
|
||||
if (not stat.S_ISREG(opened.st_mode)
|
||||
or (opened.st_dev, opened.st_ino) != (before.st_dev, before.st_ino)):
|
||||
raise ValueError(f"staged artifact 文件读取时被替换:{relative}")
|
||||
if total_bytes + opened.st_size > max_bytes:
|
||||
raise ValueError(f"staged artifact 总字节超过 {max_bytes}")
|
||||
|
||||
chunks = []
|
||||
while True:
|
||||
chunk = os.read(file_fd, 1024 * 1024)
|
||||
if not chunk:
|
||||
break
|
||||
total_bytes += len(chunk)
|
||||
if total_bytes > max_bytes:
|
||||
raise ValueError(f"staged artifact 总字节超过 {max_bytes}")
|
||||
chunks.append(chunk)
|
||||
data = b"".join(chunks)
|
||||
after = os.fstat(file_fd)
|
||||
stable_fields = ("st_dev", "st_ino", "st_size", "st_mtime_ns", "st_ctime_ns")
|
||||
if (any(getattr(opened, field) != getattr(after, field) for field in stable_fields)
|
||||
or len(data) != after.st_size):
|
||||
raise ValueError(f"staged artifact 文件读取中发生改写:{relative}")
|
||||
files[relative] = data
|
||||
finally:
|
||||
os.close(file_fd)
|
||||
|
||||
try:
|
||||
if not stat.S_ISDIR(os.fstat(root_fd).st_mode):
|
||||
raise ValueError("staged artifact 根不是目录")
|
||||
capture_dir(root_fd, "")
|
||||
finally:
|
||||
os.close(root_fd)
|
||||
|
||||
immutable_files = MappingProxyType(dict(files))
|
||||
return ArtifactSnapshot(
|
||||
files=immutable_files,
|
||||
artifact_hash=_artifact_hash(immutable_files),
|
||||
file_count=len(immutable_files),
|
||||
total_bytes=total_bytes,
|
||||
)
|
||||
|
||||
|
||||
def write_snapshot_stream(root: Path, output) -> None:
|
||||
"""把可信快照写成一行 manifest 加连续文件字节,供 Node runner 无损读取。"""
|
||||
snapshot = capture_artifact_snapshot(root)
|
||||
entries = [
|
||||
{"path": relative, "size": len(snapshot.files[relative])}
|
||||
for relative in sorted(snapshot.files, key=lambda value: value.encode("utf-8"))
|
||||
]
|
||||
manifest = {
|
||||
"schemaVersion": "artifact-snapshot-stream/1",
|
||||
"artifactHash": snapshot.artifact_hash,
|
||||
"fileCount": snapshot.file_count,
|
||||
"totalBytes": snapshot.total_bytes,
|
||||
"entries": entries,
|
||||
}
|
||||
manifest_bytes = json.dumps(
|
||||
manifest, ensure_ascii=False, sort_keys=True, separators=(",", ":"), allow_nan=False,
|
||||
).encode("utf-8")
|
||||
output.write(_SNAPSHOT_STREAM_MAGIC)
|
||||
output.write(manifest_bytes + b"\n")
|
||||
for entry in entries:
|
||||
output.write(snapshot.files[entry["path"]])
|
||||
|
||||
|
||||
def _main(argv: list[str]) -> int:
|
||||
"""仅暴露只读 stream 子命令;错误写 stderr,stdout 永远不混入诊断文本。"""
|
||||
if len(argv) != 3 or argv[1] != "--stream":
|
||||
print("用法: artifact_snapshot.py --stream <staged-root>", file=sys.stderr)
|
||||
return 2
|
||||
try:
|
||||
write_snapshot_stream(Path(argv[2]), sys.stdout.buffer)
|
||||
except Exception as exc: # noqa: BLE001 —— 子进程边界需把原始类别留给 Node 映射为 tester_error
|
||||
print(f"{type(exc).__name__}: {exc}", file=sys.stderr)
|
||||
return 2
|
||||
return 0
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
raise SystemExit(_main(sys.argv))
|
||||
477
cheap-worker/baseline_gates.py
Normal file
477
cheap-worker/baseline_gates.py
Normal file
@ -0,0 +1,477 @@
|
||||
"""baseline_gates.py — 生成线验收 v3 三批基线闸门(W-AXIS 收口 R1)。
|
||||
|
||||
把设计档 §5.2/§5.3/§5.4 的三批基线达标阈值落成机器断言,输入批结果、输出 PASS/FAIL + 逐项明细:
|
||||
|
||||
· fresh25(§5.3 波 3 新基线):五品类各五局共 25 distinct gid;accepted≥20/25 且每品类≥3/5、
|
||||
firstPassAccepted≥18/25、writer repair 启动率≤5/25、rescuedByRoll≤5/25、单局验收成本≤¥1.5、
|
||||
parentRun 全链≤¥15。分子只认 row.accepted(硬证完整 + 双 Judge 共识 + finalPostguard 全绿的最终权威),
|
||||
acceptedAfterRepair 只作修复救回率观测,绝不作成功率分子;inconclusive/tester_error 单列不剔出分母。
|
||||
· historical11(§5.2 固定预期表):假阳放行 0 + 逐局预期相符 + needs_human 项定标前不得自动 accept
|
||||
(reject/inconclusive 挂起等创始人定标,闸门不自动 PASS)。预期表 fixture =
|
||||
contracts/play-loop/historical-11-expectations.json。
|
||||
· shadow20(§5.4 生产分布 shadow):proof 完整率 100% + 确认假阳 0 + accepted 与 problems/缺证/矛盾共存 0
|
||||
+ tester_error=0(20 局口径 <5%)+ inconclusive≤1(<10%)+ 人工复核覆盖;固定六项(commit/Chrome/
|
||||
Actor+Judge 模型/prompt 版本/配置快照/人工标签)漂移即 fail;fresh25 达标为硬前置(§5.4「fresh 25 达标后」)。
|
||||
|
||||
三批依赖顺序在代码里是可选链式校验:fresh25 结果可传入 historical/shadow20 闸门——historical 收到未过的
|
||||
fresh25 时 warn 并阻断 PASS(顺序是执行建议,不阻断单跑的指标评估);shadow20 按 §5.4 硬查前置。
|
||||
本模块纯函数 + 罐头单测,零模型零 I/O(除读预期表 fixture);真跑留 mini-desktop。
|
||||
"""
|
||||
|
||||
import json
|
||||
from pathlib import Path
|
||||
|
||||
# ────────────────────────── fresh25(设计档 §5.3)──────────────────────────
|
||||
|
||||
# 五个【有 per-genre 模板】的已覆盖品类,与 hard_genre_batch.COVERED_BRIEFS 同集(硬编码保持闸门自包含)。
|
||||
FRESH25_GENRES = ("narrative", "trpg", "heritage", "puzzle", "sim-business")
|
||||
|
||||
# 阈值严格按设计档 §5.3,不放松。
|
||||
FRESH25_THRESHOLDS = {
|
||||
"totalGames": 25, # 五品类各五局共 25 distinct gid
|
||||
"gamesPerGenre": 5,
|
||||
"minAccepted": 20, # MVP 终态成功率按 accepted 计 ≥ 20/25
|
||||
"minAcceptedPerGenre": 3, # 任一品类不低于 3/5
|
||||
"minFirstPassAccepted": 18, # firstPassAccepted ≥ 18/25(最多二掷、尚未 repair 的终态)
|
||||
"maxRepairAttempted": 5, # writer repair 启动率 ≤ 5/25
|
||||
"maxRescuedByRoll": 5, # rescuedByRoll(二掷救回)≤ 5/25
|
||||
"maxRunAcceptanceCostRmb": 1.5, # 单局验收成本 ≤ ¥1.5
|
||||
"maxParentChainCostRmb": 15.0, # parentRun 全链总成本 ≤ ¥15 硬地板
|
||||
}
|
||||
|
||||
_COST_EPS = 1e-9 # 金额比较容差(与 _final_postguard_v3 的 +1e-9 口径一致)
|
||||
|
||||
|
||||
def evaluate_fresh25_gate(rows, *, thresholds=None) -> dict:
|
||||
"""fresh25 达标断言:输入批结果行(hard_genre_batch 行 schema),输出 {pass, checks, metrics, warnings}。
|
||||
|
||||
row 消费字段:gid/genre/accepted/firstPassAccepted/repairAttempted/rescuedByRoll/
|
||||
acceptanceCostRmb/parentChainCostRmb/outcome/acceptedAfterRepair。
|
||||
"""
|
||||
th = dict(FRESH25_THRESHOLDS)
|
||||
if thresholds:
|
||||
th.update(thresholds)
|
||||
rows = [r for r in (rows or []) if isinstance(r, dict)]
|
||||
checks = []
|
||||
warnings = []
|
||||
|
||||
def check(name, ok, actual, required, detail=""):
|
||||
checks.append({"name": name, "pass": bool(ok), "actual": actual, "required": required, "detail": detail})
|
||||
|
||||
# ── 样本完整性:25 distinct gid、恰好 5 品类、每品类 5 局(分母不完整 → 基线不成立,直接 fail)。
|
||||
gids = [r.get("gid") for r in rows]
|
||||
dup_gids = sorted({g for g in gids if gids.count(g) > 1 and g is not None})
|
||||
by_genre = {}
|
||||
for r in rows:
|
||||
by_genre.setdefault(r.get("genre"), []).append(r)
|
||||
genre_counts = {g: len(by_genre.get(g) or []) for g in FRESH25_GENRES}
|
||||
extra_genres = sorted(g for g in by_genre if g not in FRESH25_GENRES)
|
||||
size_ok = (len(rows) == th["totalGames"] and not dup_gids and not extra_genres
|
||||
and all(c == th["gamesPerGenre"] for c in genre_counts.values()))
|
||||
check("sampleSize", size_ok, f"{len(rows)} 局/品类分布 {genre_counts}",
|
||||
f"{th['totalGames']} 局(每品类 {th['gamesPerGenre']},gid 不重复)",
|
||||
f"重复 gid={dup_gids}" if dup_gids else (f"非覆盖品类={extra_genres}" if extra_genres else ""))
|
||||
|
||||
# ── accepted 总数(分子 = 硬证完整 + 双 Judge 共识 + finalPostguard 全绿的最终权威)。
|
||||
accepted = [r for r in rows if r.get("accepted") is True]
|
||||
check("acceptedTotal", len(accepted) >= th["minAccepted"], f"{len(accepted)}/{len(rows)}",
|
||||
f"≥ {th['minAccepted']}/{th['totalGames']}")
|
||||
|
||||
# ── 每品类 accepted ≥ 3/5。
|
||||
per_genre_accepted = {g: sum(1 for r in by_genre.get(g) or [] if r.get("accepted") is True)
|
||||
for g in FRESH25_GENRES}
|
||||
failing_genres = sorted(g for g, c in per_genre_accepted.items() if c < th["minAcceptedPerGenre"])
|
||||
check("perGenreMin", not failing_genres, per_genre_accepted,
|
||||
f"每品类 ≥ {th['minAcceptedPerGenre']}/{th['gamesPerGenre']}",
|
||||
f"未达标品类={failing_genres}" if failing_genres else "")
|
||||
|
||||
# ── firstPassAccepted ≥ 18/25(首轮直接成功 + 二掷内成功,repair 前的终态)。
|
||||
first_pass = sum(1 for r in rows if r.get("firstPassAccepted") is True)
|
||||
check("firstPassAccepted", first_pass >= th["minFirstPassAccepted"], f"{first_pass}/{len(rows)}",
|
||||
f"≥ {th['minFirstPassAccepted']}/{th['totalGames']}")
|
||||
|
||||
# ── writer repair 启动率 ≤ 5/25。
|
||||
repair_attempted = sum(1 for r in rows if r.get("repairAttempted") is True)
|
||||
check("repairRate", repair_attempted <= th["maxRepairAttempted"], f"{repair_attempted}/{len(rows)}",
|
||||
f"≤ {th['maxRepairAttempted']}/{th['totalGames']}")
|
||||
|
||||
# ── rescuedByRoll(二掷救回,取值 2)≤ 5/25;与 hard_genre_batch 主汇总口径一致。
|
||||
rescued = sum(1 for r in rows if r.get("rescuedByRoll") == 2)
|
||||
check("rescuedByRollRate", rescued <= th["maxRescuedByRoll"], f"{rescued}/{len(rows)}",
|
||||
f"≤ {th['maxRescuedByRoll']}/{th['totalGames']}")
|
||||
|
||||
# ── 单局验收成本 ≤ ¥1.5(缺成本 = 无法证明达标 → fail-closed 该检查)。
|
||||
run_costs = []
|
||||
missing_run_cost = []
|
||||
for r in rows:
|
||||
cost = r.get("acceptanceCostRmb")
|
||||
if isinstance(cost, (int, float)) and cost == cost: # 拒 NaN(cost != cost 即 NaN)
|
||||
run_costs.append((r.get("gid"), float(cost)))
|
||||
else:
|
||||
missing_run_cost.append(r.get("gid"))
|
||||
over_run = [(g, c) for g, c in run_costs if c > th["maxRunAcceptanceCostRmb"] + _COST_EPS]
|
||||
check("perRunCost", rows != [] and not over_run and not missing_run_cost,
|
||||
f"max=¥{max((c for _, c in run_costs), default=0.0):.5f}/缺 {len(missing_run_cost)} 局",
|
||||
f"每局 ≤ ¥{th['maxRunAcceptanceCostRmb']}",
|
||||
f"超限={over_run[:3]}" if over_run else (f"缺成本 gid={missing_run_cost[:3]}" if missing_run_cost else ""))
|
||||
|
||||
# ── parentRun 全链总成本 ≤ ¥15(repair 历史 + writer + 验收的不重不漏权威总成本)。
|
||||
chain_costs = []
|
||||
missing_chain_cost = []
|
||||
for r in rows:
|
||||
cost = r.get("parentChainCostRmb")
|
||||
if isinstance(cost, (int, float)) and cost == cost:
|
||||
chain_costs.append((r.get("gid"), float(cost)))
|
||||
else:
|
||||
missing_chain_cost.append(r.get("gid"))
|
||||
over_chain = [(g, c) for g, c in chain_costs if c > th["maxParentChainCostRmb"] + _COST_EPS]
|
||||
check("parentChainCost", rows != [] and not over_chain and not missing_chain_cost,
|
||||
f"max=¥{max((c for _, c in chain_costs), default=0.0):.5f}/缺 {len(missing_chain_cost)} 局",
|
||||
f"全链 ≤ ¥{th['maxParentChainCostRmb']}",
|
||||
f"超限={over_chain[:3]}" if over_chain else
|
||||
(f"缺成本 gid={missing_chain_cost[:3]}" if missing_chain_cost else ""))
|
||||
|
||||
# ── 观测项(不作闸门):四态分布 + acceptedAfterRepair(只统计修复救回率,绝不当成功率分子)。
|
||||
outcomes = {name: sum(1 for r in rows if r.get("outcome") == name)
|
||||
for name in ("accept", "reject", "inconclusive", "tester_error")}
|
||||
repaired = sum(1 for r in rows if r.get("acceptedAfterRepair") is True)
|
||||
if outcomes["inconclusive"] or outcomes["tester_error"]:
|
||||
warnings.append(f"inconclusive={outcomes['inconclusive']} tester_error={outcomes['tester_error']} "
|
||||
f"单列观测,未并入游戏失败也未剔出分母")
|
||||
metrics = {
|
||||
"accepted": len(accepted), "firstPassAccepted": first_pass,
|
||||
"repairAttempted": repair_attempted, "rescuedByRoll": rescued,
|
||||
"perGenreAccepted": per_genre_accepted, "outcomes": outcomes,
|
||||
"acceptedAfterRepair_observation": repaired, # 修复救回率观测,不是成功率分子
|
||||
"maxRunAcceptanceCostRmb": max((c for _, c in run_costs), default=None),
|
||||
"maxParentChainCostRmb": max((c for _, c in chain_costs), default=None),
|
||||
}
|
||||
return {"pass": all(c["pass"] for c in checks), "checks": checks,
|
||||
"metrics": metrics, "warnings": warnings, "thresholds": th}
|
||||
|
||||
|
||||
# ────────────────────────── historical11(设计档 §5.2)──────────────────────────
|
||||
|
||||
# 预期表 fixture:11 局固定预期(3 narrative 正例 / 5 旧假阳 / puzzle-r2 真 bug / 2 疑似假阴)。
|
||||
_HISTORICAL_EXPECTATIONS_PATH = (Path(__file__).resolve().parents[1]
|
||||
/ "contracts" / "play-loop" / "historical-11-expectations.json")
|
||||
|
||||
_EXPECTED_VALUES = ("accept", "reject", "not_accept", "needs_human")
|
||||
|
||||
|
||||
def load_historical_expectations(path=None) -> dict:
|
||||
"""读并校验 11 局固定预期表;结构非法直接抛 ValueError(闸门 fixture 必须机器可信)。"""
|
||||
p = Path(path) if path else _HISTORICAL_EXPECTATIONS_PATH
|
||||
data = json.loads(p.read_text(encoding="utf-8"))
|
||||
rows = data.get("expectations")
|
||||
if not isinstance(rows, list) or len(rows) != 11:
|
||||
raise ValueError(f"historical 预期表必须恰有 11 局,实际 {len(rows) if isinstance(rows, list) else 'N/A'}")
|
||||
gids = [r.get("gid") for r in rows if isinstance(r, dict)]
|
||||
if len(set(gids)) != 11 or any(not isinstance(g, str) or not g for g in gids):
|
||||
raise ValueError("historical 预期表 gid 必须 11 个非空且不重复")
|
||||
for r in rows:
|
||||
if r.get("expected") not in _EXPECTED_VALUES:
|
||||
raise ValueError(f"historical 预期 {r.get('gid')} expected 非法:{r.get('expected')}")
|
||||
return data
|
||||
|
||||
|
||||
def evaluate_historical_gate(replay_rows, *, expectations=None, expectations_path=None,
|
||||
fresh25_gate=None) -> dict:
|
||||
"""historical replay 闸门:假阳放行 0 + 逐局预期相符 + needs_human 定标前不自动 accept。
|
||||
|
||||
replay_rows:[{gid, outcome}],outcome ∈ accept/reject/inconclusive/tester_error(historical_replay 产物)。
|
||||
fresh25_gate:可选链式校验——提供且未过 → warn 并阻断 PASS(顺序是执行建议,不阻断指标评估);None → warn 单跑。
|
||||
"""
|
||||
data = expectations if isinstance(expectations, dict) else load_historical_expectations(expectations_path)
|
||||
exp_by_gid = {r["gid"]: r for r in data["expectations"]}
|
||||
rows = [r for r in (replay_rows or []) if isinstance(r, dict)]
|
||||
replay_by_gid = {}
|
||||
for r in rows:
|
||||
gid = r.get("gid")
|
||||
if gid in exp_by_gid and gid not in replay_by_gid:
|
||||
replay_by_gid[gid] = r
|
||||
|
||||
warnings = []
|
||||
per_game = []
|
||||
false_positive_released = [] # critical 局被 accept(已知坏例放行,闸门红线 = 0)
|
||||
mismatches = [] # 逐局预期不符
|
||||
pre_calibration_accepts = [] # needs_human 局定标前被 accept(红线)
|
||||
pending_human = [] # needs_human 局待定标(挂起,阻断自动 PASS,交创始人)
|
||||
|
||||
extra = sorted({r.get("gid") for r in rows if r.get("gid") not in exp_by_gid})
|
||||
if extra:
|
||||
warnings.append(f"重放含预期表外 gid(忽略不计):{extra}")
|
||||
|
||||
for gid, exp in exp_by_gid.items():
|
||||
expected = exp["expected"]
|
||||
row = replay_by_gid.get(gid)
|
||||
if row is None:
|
||||
mismatches.append({"gid": gid, "expected": expected, "actual": None, "reason": "重放缺该局"})
|
||||
per_game.append({"gid": gid, "expected": expected, "actual": None, "match": False})
|
||||
continue
|
||||
outcome = row.get("outcome")
|
||||
match = False
|
||||
reason = ""
|
||||
if expected == "accept":
|
||||
match = outcome == "accept"
|
||||
if not match:
|
||||
reason = "正例应保留(accept),实际未 accept——验收器可能回退"
|
||||
elif expected == "not_accept":
|
||||
if outcome == "accept":
|
||||
false_positive_released.append(gid)
|
||||
reason = "旧假阳无新硬证被 accept = 确认假阳放行"
|
||||
elif outcome in ("reject", "inconclusive"):
|
||||
match = True
|
||||
else:
|
||||
reason = f"tester_error 不构成证据({outcome}),需重跑"
|
||||
elif expected == "reject":
|
||||
if outcome == "accept":
|
||||
false_positive_released.append(gid)
|
||||
reason = "真 bug 局被 accept = 假阳放行"
|
||||
elif outcome == "reject":
|
||||
match = True
|
||||
else:
|
||||
reason = f"固定预期 reject,实际 {outcome} 不符——查验收器为何拿不到硬证"
|
||||
elif expected == "needs_human":
|
||||
if outcome == "accept":
|
||||
pre_calibration_accepts.append(gid)
|
||||
reason = "疑似假阴未定标即 accept——定标前不得自动接受"
|
||||
elif outcome in ("reject", "inconclusive", "tester_error"):
|
||||
pending_human.append(gid)
|
||||
reason = "挂起等真人真浏览器定标(不计自动 accept,也不自动 PASS)"
|
||||
else:
|
||||
reason = f"未知 outcome:{outcome}"
|
||||
if not match and gid not in pending_human:
|
||||
mismatches.append({"gid": gid, "expected": expected, "actual": outcome, "reason": reason})
|
||||
per_game.append({"gid": gid, "expected": expected, "actual": outcome,
|
||||
"match": match, "pendingHuman": gid in pending_human, "detail": reason})
|
||||
|
||||
# ── 链式校验(可选):fresh25 未过 → warn + 阻断 PASS;未提供 → warn 单跑。
|
||||
chain_blocked = False
|
||||
if fresh25_gate is None:
|
||||
warnings.append("未提供 fresh25 闸门结果:单跑模式(设计档建议顺序 fresh25 → historical11 → shadow20)")
|
||||
elif not isinstance(fresh25_gate, dict) or fresh25_gate.get("pass") is not True:
|
||||
chain_blocked = True
|
||||
warnings.append("链式校验:fresh25 未达标 → historical 闸门 PASS 受阻(执行顺序建议,不阻断指标评估)")
|
||||
|
||||
checks = [
|
||||
{"name": "coverage", "pass": len(replay_by_gid) == 11,
|
||||
"actual": f"{len(replay_by_gid)}/11", "required": "11 局全部重放",
|
||||
"detail": ",".join(m["gid"] for m in mismatches if m["actual"] is None)},
|
||||
{"name": "falsePositiveRelease", "pass": not false_positive_released,
|
||||
"actual": len(false_positive_released), "required": "0(已知坏例放行数为 0)",
|
||||
"detail": ",".join(false_positive_released)},
|
||||
{"name": "preCalibrationAccept", "pass": not pre_calibration_accepts,
|
||||
"actual": len(pre_calibration_accepts), "required": "0(needs_human 定标前不得 accept)",
|
||||
"detail": ",".join(pre_calibration_accepts)},
|
||||
{"name": "perGameExpectation", "pass": not mismatches,
|
||||
"actual": f"{11 - len(mismatches)}/11 相符", "required": "逐局固定预期相符",
|
||||
"detail": ";".join(f"{m['gid']}:{m['reason']}" for m in mismatches if m["actual"] is not None)},
|
||||
{"name": "humanCalibrationSettled", "pass": not pending_human,
|
||||
"actual": f"{len(pending_human)} 局待定标", "required": "0(疑似假阴须先真人定标)",
|
||||
"detail": ",".join(pending_human)},
|
||||
{"name": "chainFresh25", "pass": not chain_blocked,
|
||||
"actual": "受阻" if chain_blocked else "通过/单跑", "required": "fresh25 达标(提供时)"},
|
||||
]
|
||||
return {"pass": all(c["pass"] for c in checks), "checks": checks, "perGame": per_game,
|
||||
"blockedOnHumanCalibration": pending_human,
|
||||
"falsePositiveReleased": false_positive_released,
|
||||
"warnings": warnings}
|
||||
|
||||
|
||||
# ────────────────────────── shadow20(设计档 §5.4)──────────────────────────
|
||||
|
||||
SHADOW20_MIN_SAMPLES = 20 # 「连续不少于 20 个真实生产 prompt」
|
||||
|
||||
# 固定六项:代码 commit / Chrome 版本 / Actor 模型 / Judge 模型 / prompt 版本 / 配置快照(+ 人工标签占位)。
|
||||
SHADOW20_IDENTITY_FIELDS = ("commitHash", "chromeVersion", "actorModel", "judgeModel",
|
||||
"promptVersion", "configSnapshotHash")
|
||||
|
||||
# 20 局口径下 <5% → tester_error 必须 0;<10% → inconclusive 最多 1(设计档 §5.4 明文换算)。
|
||||
SHADOW20_THRESHOLDS = {
|
||||
"minSamples": SHADOW20_MIN_SAMPLES,
|
||||
"maxConfirmedFalsePositives": 0,
|
||||
"maxAcceptedWithProblems": 0,
|
||||
"maxTesterError": 0,
|
||||
"maxInconclusive": 1,
|
||||
"minOrdinaryAcceptHumanReviewed": 5, # 普通 accept 至少抽 5(accept 不足 5 则全查)
|
||||
}
|
||||
|
||||
|
||||
def build_shadow_plan(prompts, *, commit_hash, chrome_version, actor_model, judge_model,
|
||||
prompt_version, config_snapshot_hash, human_labels=None, note="") -> dict:
|
||||
"""shadow20 跑批计划:固定六项 + ≥20 连续生产 prompt + 人工标签占位。任一固定项缺失/样本不足 → ValueError。"""
|
||||
prompts = list(prompts or [])
|
||||
if len(prompts) < SHADOW20_MIN_SAMPLES:
|
||||
raise ValueError(f"shadow20 需连续 ≥{SHADOW20_MIN_SAMPLES} 个真实生产 prompt,实际 {len(prompts)}")
|
||||
identity = {"commitHash": commit_hash, "chromeVersion": chrome_version,
|
||||
"actorModel": actor_model, "judgeModel": judge_model,
|
||||
"promptVersion": prompt_version, "configSnapshotHash": config_snapshot_hash}
|
||||
empty = [k for k, v in identity.items() if not (isinstance(v, str) and v.strip())]
|
||||
if empty:
|
||||
raise ValueError(f"shadow20 固定六项不得为空:{empty}")
|
||||
return {
|
||||
"schemaVersion": "shadow20-plan/1",
|
||||
"identity": identity,
|
||||
"prompts": prompts,
|
||||
# 人工标签占位:真跑后由人工逐局回填(分歧/reject/rescued/inconclusive/tester_error 必查,accept 抽 ≥5)。
|
||||
"humanLabels": dict(human_labels or {}),
|
||||
"note": str(note or ""),
|
||||
"acceptanceMode": "v3_shadow", # shadow 新生产 run 只跑 v3、内测隔离、冻结自动发布(§3.10)
|
||||
}
|
||||
|
||||
|
||||
async def run_shadow_batch(plan, *, run_one) -> list:
|
||||
"""shadow runner 框架:串行编排 ≥20 局独立 v3_shadow,每局盖固定身份戳。
|
||||
|
||||
run_one:注入的真跑 callable(async/sync 均可)—— (prompt, identity) → 单局结果行(mini-desktop 提供真实现,
|
||||
经 cheap_verify.run_acceptance_v3(mode=v3_shadow) 产出行);本地单测注入罐头函数,不烧真模型。
|
||||
本框架只负责串行调度 + 身份固化 + 行规范化,不碰模型。
|
||||
"""
|
||||
if not isinstance(plan, dict) or not isinstance(plan.get("prompts"), list):
|
||||
raise ValueError("shadow plan 非法")
|
||||
identity = plan.get("identity") if isinstance(plan.get("identity"), dict) else {}
|
||||
rows = []
|
||||
for index, prompt in enumerate(plan["prompts"]):
|
||||
result = run_one(prompt, dict(identity))
|
||||
if hasattr(result, "__await__"):
|
||||
result = await result
|
||||
row = dict(result) if isinstance(result, dict) else {"outcome": "tester_error", "raw": result}
|
||||
# 每局固化身份 + 序号,供闸门核对固定六项零漂移;人工标签占位随行带出待回填。
|
||||
row.setdefault("gid", f"shadow20-{index + 1:02d}")
|
||||
row["identity"] = dict(identity)
|
||||
row["planIndex"] = index
|
||||
rows.append(row)
|
||||
return rows
|
||||
|
||||
|
||||
def _identity_drift(rows, plan) -> list:
|
||||
"""核对所有行固定六项逐字一致,且与 plan 一致(漂移即 fail——shadow 的可比性前提)。"""
|
||||
drifts = []
|
||||
reference = None
|
||||
if isinstance(plan, dict) and isinstance(plan.get("identity"), dict):
|
||||
reference = {k: plan["identity"].get(k) for k in SHADOW20_IDENTITY_FIELDS}
|
||||
for row in rows:
|
||||
identity = row.get("identity") if isinstance(row.get("identity"), dict) else {}
|
||||
values = {k: identity.get(k) for k in SHADOW20_IDENTITY_FIELDS}
|
||||
if any(not (isinstance(v, str) and v) for v in values.values()):
|
||||
drifts.append({"gid": row.get("gid"), "reason": "固定六项有空值"})
|
||||
continue
|
||||
if reference is None:
|
||||
reference = values
|
||||
elif values != reference:
|
||||
diff = [k for k in SHADOW20_IDENTITY_FIELDS if values.get(k) != reference.get(k)]
|
||||
drifts.append({"gid": row.get("gid"), "reason": f"身份漂移:{diff}"})
|
||||
return drifts
|
||||
|
||||
|
||||
def evaluate_shadow20_gate(rows, *, fresh25_gate=None, plan=None, thresholds=None) -> dict:
|
||||
"""shadow20 达标断言(§5.4):proof 完整率 100% + 确认假阳 0 + accepted 无 problems/缺证/矛盾共存
|
||||
+ tester_error=0 + inconclusive≤1 + 人工复核覆盖 + 固定六项零漂移;fresh25 达标为硬前置。
|
||||
|
||||
row 消费字段:gid/outcome/proofComplete/confirmedFalsePositive/problems/contradictions/proofMissing/
|
||||
rescued/discrepancy/humanReviewed/identity{六项}。
|
||||
"""
|
||||
th = dict(SHADOW20_THRESHOLDS)
|
||||
if thresholds:
|
||||
th.update(thresholds)
|
||||
rows = [r for r in (rows or []) if isinstance(r, dict)]
|
||||
warnings = []
|
||||
checks = []
|
||||
|
||||
def check(name, ok, actual, required, detail=""):
|
||||
checks.append({"name": name, "pass": bool(ok), "actual": actual, "required": required, "detail": detail})
|
||||
|
||||
# ── 硬前置(§5.4「fresh 25 达标后」):未提供 → 前置无法验证,阻断 PASS(warn);提供且未过 → 硬 fail。
|
||||
if fresh25_gate is None:
|
||||
warnings.append("未提供 fresh25 闸门结果:shadow20 硬前置无法验证,PASS 阻断(§5.4 顺序硬要求)")
|
||||
prereq_ok, prereq_actual = False, "未验证"
|
||||
elif isinstance(fresh25_gate, dict) and fresh25_gate.get("pass") is True:
|
||||
prereq_ok, prereq_actual = True, "fresh25 达标"
|
||||
else:
|
||||
prereq_ok, prereq_actual = False, "fresh25 未达标"
|
||||
check("prerequisiteFresh25", prereq_ok, prereq_actual, "fresh25 闸门通过(§5.4 硬前置)")
|
||||
|
||||
check("sampleSize", len(rows) >= th["minSamples"], f"{len(rows)} 局",
|
||||
f"连续 ≥ {th['minSamples']} 个真实生产 prompt")
|
||||
|
||||
drifts = _identity_drift(rows, plan)
|
||||
check("identityFixed", rows != [] and not drifts, f"{len(rows) - len(drifts)}/{len(rows)} 行身份一致",
|
||||
"固定六项(commit/Chrome/Actor/Judge/prompt/配置)逐字一致",
|
||||
";".join(f"{d['gid']}:{d['reason']}" for d in drifts[:3]))
|
||||
|
||||
accepted = [r for r in rows if r.get("outcome") == "accept"]
|
||||
tester_errors = [r for r in rows if r.get("outcome") == "tester_error"]
|
||||
inconclusives = [r for r in rows if r.get("outcome") == "inconclusive"]
|
||||
|
||||
# ── accepted 的 proof 完整率 100%(缺 proofComplete 字段 = 无法证明完整 → fail-closed)。
|
||||
proof_incomplete = [r.get("gid") for r in accepted if r.get("proofComplete") is not True]
|
||||
check("proofCompleteOnAccepted", not proof_incomplete,
|
||||
f"{len(accepted) - len(proof_incomplete)}/{len(accepted)} 完整", "accepted proof 完整率 100%",
|
||||
f"不完整={proof_incomplete[:3]}")
|
||||
|
||||
# ── accepted 与 problems/缺证/矛盾共存 0。
|
||||
conflicted = [r.get("gid") for r in accepted
|
||||
if r.get("problems") or r.get("contradictions") or r.get("proofMissing") is True]
|
||||
check("acceptedWithoutConflict", not conflicted,
|
||||
f"{len(accepted) - len(conflicted)}/{len(accepted)} 干净", "accepted 与 problems/缺证/矛盾共存 0",
|
||||
f"共存={conflicted[:3]}")
|
||||
|
||||
confirmed_fp = [r.get("gid") for r in rows if r.get("confirmedFalsePositive") is True]
|
||||
check("confirmedFalsePositives", not confirmed_fp, len(confirmed_fp), "确认假阳 0", ",".join(confirmed_fp))
|
||||
|
||||
check("testerError", len(tester_errors) <= th["maxTesterError"], f"{len(tester_errors)}/{len(rows)}",
|
||||
f"≤ {th['maxTesterError']}(20 局口径 <5% → 0)", ",".join(r.get("gid", "?") for r in tester_errors[:3]))
|
||||
|
||||
check("inconclusive", len(inconclusives) <= th["maxInconclusive"], f"{len(inconclusives)}/{len(rows)}",
|
||||
f"≤ {th['maxInconclusive']}(20 局口径 <10%)", ",".join(r.get("gid", "?") for r in inconclusives[:3]))
|
||||
|
||||
# ── 人工复核覆盖:全部 reject/rescued/inconclusive/tester_error/分歧必查 + 普通 accept 至少抽 5。
|
||||
must_review = [r for r in rows
|
||||
if r.get("outcome") in ("reject", "inconclusive", "tester_error")
|
||||
or r.get("rescued") is True or r.get("discrepancy") is True]
|
||||
unreviewed = [r.get("gid") for r in must_review if r.get("humanReviewed") is not True]
|
||||
reviewed_accepts = sum(1 for r in accepted if r.get("humanReviewed") is True)
|
||||
accept_quota = min(th["minOrdinaryAcceptHumanReviewed"], len(accepted))
|
||||
review_ok = not unreviewed and reviewed_accepts >= accept_quota
|
||||
check("humanReviewCoverage", review_ok,
|
||||
f"必查 {len(must_review) - len(unreviewed)}/{len(must_review)},accept 抽查 {reviewed_accepts}/{accept_quota}",
|
||||
f"分歧/reject/rescued/inconclusive/tester_error 全查 + 普通 accept ≥ {accept_quota}",
|
||||
f"未复核={unreviewed[:3]}")
|
||||
|
||||
outcomes = {"accept": len(accepted), "reject": sum(1 for r in rows if r.get("outcome") == "reject"),
|
||||
"inconclusive": len(inconclusives), "tester_error": len(tester_errors)}
|
||||
return {"pass": all(c["pass"] for c in checks), "checks": checks,
|
||||
"metrics": {"outcomes": outcomes, "sampleSize": len(rows),
|
||||
"proofCompleteOnAccepted": len(accepted) - len(proof_incomplete)},
|
||||
"warnings": warnings, "thresholds": th}
|
||||
|
||||
|
||||
# ────────────────────────── CLI(mini-desktop 真跑后对账用)──────────────────────────
|
||||
|
||||
def _cli() -> int:
|
||||
import argparse
|
||||
parser = argparse.ArgumentParser(description="三批基线闸门(fresh25 / historical / shadow20)")
|
||||
parser.add_argument("gate", choices=("fresh25", "historical", "shadow20"))
|
||||
parser.add_argument("rows_json", help="批结果行 JSON 数组文件")
|
||||
parser.add_argument("--fresh25-result", help="shadow20/historical 链式校验:fresh25 闸门结果 JSON 文件")
|
||||
parser.add_argument("--expectations", help="historical 预期表(默认 contracts/play-loop/historical-11-expectations.json)")
|
||||
parser.add_argument("--plan", help="shadow20 计划 JSON(核对固定六项)")
|
||||
args = parser.parse_args()
|
||||
rows = json.loads(Path(args.rows_json).read_text(encoding="utf-8"))
|
||||
fresh25_gate = None
|
||||
if args.fresh25_result:
|
||||
fresh25_gate = json.loads(Path(args.fresh25_result).read_text(encoding="utf-8"))
|
||||
if args.gate == "fresh25":
|
||||
report = evaluate_fresh25_gate(rows)
|
||||
elif args.gate == "historical":
|
||||
report = evaluate_historical_gate(rows, expectations_path=args.expectations, fresh25_gate=fresh25_gate)
|
||||
else:
|
||||
plan = json.loads(Path(args.plan).read_text(encoding="utf-8")) if args.plan else None
|
||||
report = evaluate_shadow20_gate(rows, fresh25_gate=fresh25_gate, plan=plan)
|
||||
print(json.dumps(report, ensure_ascii=False, indent=2))
|
||||
return 0 if report.get("pass") else 1
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
raise SystemExit(_cli())
|
||||
@ -1,7 +1,8 @@
|
||||
"""cheap_service_app.py — 便宜档 · 独立 Agent Service 服务壳(镜像 tier2 service/app.py,cheap 各起进程;决策②)。
|
||||
|
||||
把便宜档从「裸 HTTP /generate + 进程内 for-resume」归并到 AgentScope Service /chat:六工具/续修/软预算/trace
|
||||
经工厂 per-turn 注入,续修由阶段一① 的 RepairMiddleware 在 finish 点拦九门续跑(取代进程内 resume 循环)。
|
||||
把便宜档从「裸 HTTP /generate + 进程内 for-resume」归并到 AgentScope Service /chat:七工具/软预算/trace
|
||||
经工厂 per-turn 注入。v3 模式不装旧 RepairMiddleware;writer 完成后由 driver 跑机械门与 v3,只有 final verified
|
||||
reject 才允许同一 session 修一次。RepairMiddleware 仅保留给显式 v1/v2 历史回放。
|
||||
与 tier2 Service 各起独立进程:cheap credential 走 OpenAI 兼容路(base+/v1)、预算 soft+¥10 软目标(成本上界靠
|
||||
max_repairs 优雅终止、三闸=150 失控兜底)、六工具面、并发按 session 从进程内端口池派生(避撞固定 4320/9222)。
|
||||
两工厂经会话注册表 sidecar 把框架分配的 session_id 解析回后端 gameId(C2:产物目录/评门/回调统一后端 gameId)。
|
||||
@ -12,8 +13,11 @@ max_repairs 优雅终止、三闸=150 失控兜底)、六工具面、并发按 s
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import hashlib
|
||||
import json
|
||||
import os
|
||||
import queue as _queue
|
||||
import stat
|
||||
import sys
|
||||
import threading
|
||||
from pathlib import Path
|
||||
@ -28,6 +32,13 @@ if TYPE_CHECKING: # pragma: no cover
|
||||
|
||||
# 服务标题(OpenAPI docs 显示)。
|
||||
SERVICE_TITLE = "cheap-littlejs-agent-service"
|
||||
_SESSION_CFG_MAX_BYTES = 1 * 1024 * 1024
|
||||
_REFERENCE_RECEIPTS_FILE = ".reference-receipts.json"
|
||||
|
||||
|
||||
def _reference_snapshot_dir(session_cfg_path: Path) -> Path:
|
||||
"""由固定 sidecar 路径派生 session 快照目录,不读取配置中的路径。"""
|
||||
return session_cfg_path.with_name(f"{session_cfg_path.stem}.reference-assets")
|
||||
|
||||
# ── 并发端口池(决策③):便宜档并发≤15,play 用固定 4320/9222 会撞;按 session 从池派生唯一端口对 ──
|
||||
# 与 tier2(4330/9322)、旧 cheap 串行默认(4320/9222)错开,避免同机多 Service 撞端口。池大小 16 > 上限 15。
|
||||
@ -75,21 +86,191 @@ def _read_session_cfg(session_id: str) -> dict:
|
||||
因 AgentScope 工厂签名固定 (user_id, agent_id, session_id)、拿不到后端 gameId 也拿不到 session 记录,
|
||||
且 AgentData/SessionConfig 无自由字段,改用 worker↔Service 同机共享 FS sidecar
|
||||
game-runtime/games/_cheap-sessions/<session_id>.json 把「session_id → 后端 gameId + 本 session 配置」传给工厂。
|
||||
缺失 / 坏 JSON → {}(工厂据此回落 game_id=session_id;正常路 driver 恒在 /chat 前写、不会缺;best-effort,绝不抛)。
|
||||
缺失 sidecar → {},保持无 sidecar 的旧路径兼容;已存在的文件若 JSON 损坏、顶层不是对象,或派生
|
||||
snapshot 存在但缺少冻结 policy,均直接拒绝,避免在工具工厂边界回落到活目录。
|
||||
存在的文件必须是固定上限内的普通文件,symlink、特殊文件、超限或读取竞态直接拒绝。
|
||||
"""
|
||||
import json # noqa: PLC0415
|
||||
|
||||
import cheap_run # noqa: PLC0415
|
||||
|
||||
p = cheap_run.session_cfg_path(session_id)
|
||||
if not p.exists():
|
||||
return {}
|
||||
try:
|
||||
obj = json.loads(p.read_text(encoding="utf-8"))
|
||||
return obj if isinstance(obj, dict) else {}
|
||||
except Exception as e: # noqa: BLE001 —— 坏 sidecar 回落 session_id,不阻断
|
||||
print(f"[cheap-service] session-cfg 读失败(回落 game_id=session_id):{type(e).__name__}: {e}", flush=True)
|
||||
return {}
|
||||
before = os.lstat(p)
|
||||
except FileNotFoundError:
|
||||
# 只有 cfg 与派生 snapshot 都不存在才是真正缺 sidecar;孤立冻结 snapshot 必须拒绝 live 回落。
|
||||
snapshot_dir = _reference_snapshot_dir(p)
|
||||
try:
|
||||
os.lstat(snapshot_dir)
|
||||
except FileNotFoundError:
|
||||
return {}
|
||||
except OSError as e:
|
||||
raise ValueError("reference asset session snapshot 状态不可读") from e
|
||||
raise ValueError("reference asset session snapshot 存在但 session-cfg 缺失")
|
||||
if stat.S_ISLNK(before.st_mode) or not stat.S_ISREG(before.st_mode):
|
||||
raise ValueError("session-cfg 必须是普通文件且不得为 symlink")
|
||||
if before.st_size > _SESSION_CFG_MAX_BYTES:
|
||||
raise ValueError("session-cfg 超过固定读取上限")
|
||||
try:
|
||||
flags = os.O_RDONLY | getattr(os, "O_NOFOLLOW", 0)
|
||||
fd = os.open(p, flags)
|
||||
try:
|
||||
opened = os.fstat(fd)
|
||||
if (opened.st_dev, opened.st_ino) != (before.st_dev, before.st_ino):
|
||||
raise ValueError("session-cfg 读取期间发生替换")
|
||||
chunks = []
|
||||
total = 0
|
||||
while True:
|
||||
chunk = os.read(fd, min(64 * 1024, _SESSION_CFG_MAX_BYTES + 1 - total))
|
||||
if not chunk:
|
||||
break
|
||||
total += len(chunk)
|
||||
if total > _SESSION_CFG_MAX_BYTES:
|
||||
raise ValueError("session-cfg 超过固定读取上限")
|
||||
chunks.append(chunk)
|
||||
after = os.fstat(fd)
|
||||
if any(getattr(opened, field) != getattr(after, field)
|
||||
for field in ("st_dev", "st_ino", "st_size", "st_mtime_ns", "st_ctime_ns")):
|
||||
raise ValueError("session-cfg 读取期间发生漂移")
|
||||
finally:
|
||||
os.close(fd)
|
||||
raw = b"".join(chunks)
|
||||
except OSError as e:
|
||||
raise ValueError(f"session-cfg 普通文件边界拒绝:{type(e).__name__}") from e
|
||||
except ValueError:
|
||||
raise
|
||||
|
||||
try:
|
||||
obj = json.loads(raw.decode("utf-8"))
|
||||
except (UnicodeDecodeError, json.JSONDecodeError) as e:
|
||||
print(f"[cheap-service] session-cfg JSON 非法,拒绝回落活目录:{type(e).__name__}: {e}", flush=True)
|
||||
raise ValueError("session-cfg JSON 非法") from e
|
||||
if not isinstance(obj, dict):
|
||||
raise ValueError("session-cfg 顶层必须是对象")
|
||||
|
||||
# 冻结 snapshot 与 policy 必须成对存在;只剩 snapshot 时拒绝把工具绑定回原资产活目录。
|
||||
if obj.get("reference_asset_policy") is None:
|
||||
snapshot_dir = _reference_snapshot_dir(p)
|
||||
try:
|
||||
os.lstat(snapshot_dir)
|
||||
except FileNotFoundError:
|
||||
pass
|
||||
except OSError as e:
|
||||
raise ValueError("reference asset session snapshot 状态不可读") from e
|
||||
else:
|
||||
raise ValueError("reference asset session snapshot 存在但 sidecar 缺 policy")
|
||||
return obj
|
||||
|
||||
|
||||
def _is_sha256(value) -> bool:
|
||||
"""校验 sidecar 中稳定 SHA-256 文本,拒绝宽松大小写和非字符串。"""
|
||||
return isinstance(value, str) and len(value) == 64 and all(char in "0123456789abcdef" for char in value)
|
||||
|
||||
|
||||
def _load_reference_snapshot(session_id: str, cfg: dict):
|
||||
"""按 sidecar 索引重验 session snapshot,返回 Toolkit 使用的只读 files/roots。
|
||||
|
||||
只有 cfg 显式声明冻结 policy 才进入本路径;任何目录、文件、索引、receipt 或 canonical hash 漂移
|
||||
都在 ``CheapSession`` 与工具创建前抛错,绝不读取原资产活目录或 best-effort 回落。
|
||||
"""
|
||||
policy = cfg.get("reference_asset_policy")
|
||||
if policy is None:
|
||||
return None, None
|
||||
expected_keys = {"policy_id", "mode", "snapshot_hash", "receipt_hash", "roots", "files"}
|
||||
if not isinstance(policy, dict) or set(policy) != expected_keys:
|
||||
raise ValueError("reference asset policy sidecar 结构非法")
|
||||
if policy["policy_id"] != "survivor-gold-v1" or policy["mode"] != "frozen_preflight":
|
||||
raise ValueError("reference asset policy/mode 不受信")
|
||||
if not _is_sha256(policy["snapshot_hash"]) or not _is_sha256(policy["receipt_hash"]):
|
||||
raise ValueError("reference asset sidecar hash 非法")
|
||||
if not isinstance(policy["roots"], dict) or not policy["roots"]:
|
||||
raise ValueError("reference asset roots 索引非法")
|
||||
entries = policy["files"]
|
||||
if not isinstance(entries, list) or not entries or len(entries) > 512:
|
||||
raise ValueError("reference asset 文件索引非法")
|
||||
|
||||
import artifact_snapshot # noqa: PLC0415 复用 fd/O_NOFOLLOW 可信读取边界
|
||||
import cheap_run # noqa: PLC0415
|
||||
import reference_asset_gate # noqa: PLC0415
|
||||
|
||||
snapshot_dir = _reference_snapshot_dir(cheap_run.session_cfg_path(session_id))
|
||||
try:
|
||||
root_stat = os.lstat(snapshot_dir)
|
||||
except FileNotFoundError as exc:
|
||||
raise ValueError("reference asset session snapshot 目录缺失") from exc
|
||||
if stat.S_ISLNK(root_stat.st_mode) or not stat.S_ISDIR(root_stat.st_mode):
|
||||
raise ValueError("reference asset session snapshot 根不是普通目录")
|
||||
|
||||
paths = []
|
||||
by_path = {}
|
||||
total_declared = 0
|
||||
for entry in entries:
|
||||
if not isinstance(entry, dict) or set(entry) != {"path", "size", "sha256"}:
|
||||
raise ValueError("reference asset 文件索引条目非法")
|
||||
path, size, digest = entry["path"], entry["size"], entry["sha256"]
|
||||
if (not isinstance(path, str) or not path or path in by_path
|
||||
or not isinstance(size, int) or isinstance(size, bool) or size < 0
|
||||
or not _is_sha256(digest)):
|
||||
raise ValueError("reference asset 文件索引漂移")
|
||||
paths.append(path)
|
||||
by_path[path] = entry
|
||||
total_declared += size
|
||||
if total_declared > reference_asset_gate.MAX_TOTAL_BYTES:
|
||||
raise ValueError("reference asset snapshot 超过 128 MiB 硬帽")
|
||||
if paths != sorted(paths, key=lambda item: item.encode("utf-8")):
|
||||
raise ValueError("reference asset 文件索引顺序漂移")
|
||||
|
||||
files = {}
|
||||
total_observed = 0
|
||||
for path in paths:
|
||||
try:
|
||||
captured = artifact_snapshot.capture_selected_files(
|
||||
snapshot_dir,
|
||||
[path],
|
||||
limits={
|
||||
"max_files": 1,
|
||||
"max_file_bytes": reference_asset_gate.MAX_FILE_BYTES,
|
||||
"max_record_bytes": reference_asset_gate.MAX_FILE_BYTES,
|
||||
"max_total_bytes": reference_asset_gate.MAX_FILE_BYTES,
|
||||
},
|
||||
)
|
||||
except artifact_snapshot.ArtifactSnapshotError as exc:
|
||||
raise ValueError(f"reference asset snapshot 文件拒绝:{exc.code}") from exc
|
||||
content = captured.files[path]
|
||||
total_observed += len(content)
|
||||
expected = by_path[path]
|
||||
if len(content) != expected["size"] or hashlib.sha256(content).hexdigest() != expected["sha256"]:
|
||||
raise ValueError("reference asset 文件索引漂移")
|
||||
if total_observed > reference_asset_gate.MAX_TOTAL_BYTES:
|
||||
raise ValueError("reference asset snapshot 超过 128 MiB 硬帽")
|
||||
files[path] = content
|
||||
if total_observed != total_declared:
|
||||
raise ValueError("reference asset 文件索引总量漂移")
|
||||
if reference_asset_gate.snapshot_hash(files) != policy["snapshot_hash"]:
|
||||
raise ValueError("reference asset snapshot hash 漂移")
|
||||
|
||||
try:
|
||||
receipt_capture = artifact_snapshot.capture_selected_files(
|
||||
snapshot_dir,
|
||||
[_REFERENCE_RECEIPTS_FILE],
|
||||
limits={"max_files": 1, "max_file_bytes": 1 * 1024 * 1024,
|
||||
"max_record_bytes": 1 * 1024 * 1024, "max_total_bytes": 1 * 1024 * 1024},
|
||||
)
|
||||
except artifact_snapshot.ArtifactSnapshotError as exc:
|
||||
raise ValueError(f"reference asset receipt 拒绝:{exc.code}") from exc
|
||||
receipt_bytes = receipt_capture.files[_REFERENCE_RECEIPTS_FILE]
|
||||
if hashlib.sha256(receipt_bytes).hexdigest() != policy["receipt_hash"]:
|
||||
raise ValueError("reference asset receipt hash 漂移")
|
||||
try:
|
||||
receipt_payload = json.loads(receipt_bytes.decode("utf-8"))
|
||||
except Exception as exc: # noqa: BLE001 receipt 必须是 canonical JSON 对象
|
||||
raise ValueError("reference asset receipt 内容非法") from exc
|
||||
receipts = receipt_payload.get("receipts") if isinstance(receipt_payload, dict) else None
|
||||
if (not isinstance(receipts, list) or not receipts
|
||||
or any(not isinstance(item, dict)
|
||||
or item.get("finalSnapshotHash") != policy["snapshot_hash"] for item in receipts)):
|
||||
raise ValueError("reference asset receipt 与 snapshot 不一致")
|
||||
if reference_asset_gate.canonical_json_bytes(receipt_payload) != receipt_bytes:
|
||||
raise ValueError("reference asset receipt canonical 字节漂移")
|
||||
return files, policy["roots"]
|
||||
|
||||
|
||||
def _resolve_external_game_id(session_id: str, cfg: dict) -> str:
|
||||
@ -135,7 +316,12 @@ async def _cheap_tools_factory(user_id: str, agent_id: str, session_id: str) ->
|
||||
cfg = _read_session_cfg(session_id)
|
||||
game_id = _resolve_external_game_id(session_id, cfg) # C2:session_id → 后端 gameId
|
||||
write_whitelist = _resolve_write_whitelist(cfg) # I1:create None / restricted 缺白名单则空集(禁写)
|
||||
session = CheapSession(game_id=game_id) # 六工具据后端 gameId 管 amgen-<后端 gameId> 目录(与 scaffold/prompt 一致)
|
||||
reference_files, reference_roots = _load_reference_snapshot(session_id, cfg)
|
||||
session = CheapSession(
|
||||
game_id=game_id,
|
||||
reference_files=reference_files,
|
||||
reference_roots=reference_roots,
|
||||
) # 六工具据后端 gameId 管产物目录;显式 policy 同时绑定已复核的 session 只读快照。
|
||||
toolkit = build_toolkit(session, write_whitelist=write_whitelist)
|
||||
return _extract_function_tools(toolkit)
|
||||
|
||||
@ -292,15 +478,14 @@ _COLLECTOR_CLS = None
|
||||
|
||||
|
||||
async def _cheap_middlewares_factory(user_id: str, agent_id: str, session_id: str) -> list:
|
||||
"""extra_agent_middlewares 工厂:每回合产 [collector, trace, 续修, 软预算熔断] 四件(镜像 tier2,cheap 独立参数)。
|
||||
"""按验收模式装配中间件;v3 为 collector/trace/breaker,历史模式才额外装旧 repair。
|
||||
|
||||
续修 check = finish 点独立跑便宜档九门(cheap_gates.run_cheap_gates 经 to_thread,端口按 session 从池派生)→
|
||||
judge_cheap_verdict;异常兜底放工厂闭包 try(T2-b:镜像 tier2 app.py:210-216,不下沉 RepairMiddleware 共用类)。
|
||||
breaker soft_budget=True(决策①:¥10 转软目标、超预算软停交尽力产物);成本上界靠 max_repairs=6 优雅终止,
|
||||
历史续修 check = finish 点独立跑便宜档九门(cheap_gates.run_cheap_gates 经 to_thread,端口按 session 从池派生)→
|
||||
judge_cheap_verdict;v3 不调用此闭包,机械门改由 driver 在回合结束后只跑一次。
|
||||
breaker soft_budget=True;v3 只有一次协议修复,max_repairs=6 只约束历史回放;
|
||||
三道次数/轮数闸(max_tool_calls/max_model_calls/max_iters)抬到 150 当纯失控兜底(Codex C4:不设 max_tool_calls
|
||||
会吃默认 60 先撞);超时按最坏九门 ~390s 放宽(C3)。评门 game_id = 后端 gameId(经 sidecar 解析,C2)。
|
||||
注入序(外→内):collector / tracer / repair / breaker;框架在其外另前置 InboxMiddleware(on_reasoning 每轮
|
||||
drain inbox 后透传全部 evt、不吞 finish),故 on_reasoning 链 = [Inbox, tracer, repair],repair 仍最内层能拦原始 finish。
|
||||
注入序(外→内):collector / tracer / [legacy repair] / breaker;turns 观测件最后追加且不改控制流。
|
||||
"""
|
||||
import asyncio # noqa: PLC0415
|
||||
|
||||
@ -315,7 +500,7 @@ async def _cheap_middlewares_factory(user_id: str, agent_id: str, session_id: st
|
||||
)
|
||||
|
||||
# C2:评门/collector/trace 全绑后端 gameId(经会话注册表 sidecar 把 session_id 解析回后端 gameId;与 driver
|
||||
# scaffold、system prompt 的 ⟦G⟧、六工具写目录一致)。缺映射则回落 session_id(响亮失败,见 _resolve_external_game_id)。
|
||||
# scaffold、system prompt 的 ⟦G⟧、七工具写目录一致)。缺映射则回落 session_id(响亮失败,见 _resolve_external_game_id)。
|
||||
_cfg = _read_session_cfg(session_id)
|
||||
game_id = _resolve_external_game_id(session_id, _cfg)
|
||||
# 面四断点③:从 sidecar 取 worker 这跳写下的 W3C traceparent(driver 恒在 /chat 前写),组入站 carrier
|
||||
@ -392,21 +577,27 @@ async def _cheap_middlewares_factory(user_id: str, agent_id: str, session_id: st
|
||||
return judge_cheap_verdict(verdict, game_id=game_id,
|
||||
staged_dir=cheap_run.wg1_game_dir(game_id))
|
||||
|
||||
repair = RepairMiddleware(
|
||||
check=_cheap_check,
|
||||
max_repairs=max_repairs,
|
||||
# 实测已花 ¥ 超上限(非预估软停标记):配① 的 repairs>0 保护,首个未绿 finish 必先修一次再因预算放行。
|
||||
budget_exhausted=lambda: (
|
||||
breaker._rmb_gate_active and breaker.spent_rmb >= breaker.rmb_hard_limit),
|
||||
)
|
||||
acceptance_mode = str(genconfig.get("acceptance", "mode", "v3_shadow"))
|
||||
repair = None
|
||||
if acceptance_mode not in ("v3", "v3_shadow"):
|
||||
# 仅历史 v1/v2 回放保留九门 gameplay resume;v3 的唯一修复权归 final verified reject。
|
||||
repair = RepairMiddleware(
|
||||
check=_cheap_check,
|
||||
max_repairs=max_repairs,
|
||||
budget_exhausted=lambda: (
|
||||
breaker._rmb_gate_active and breaker.spent_rmb >= breaker.rmb_hard_limit),
|
||||
)
|
||||
|
||||
# 收口采集(cost/trace/repairs 跨进程回收):reply 收尾写 evidence/service-run-summary.json,供 T3 driver 组 result-out。
|
||||
collector = _get_collector_cls()(game_id=game_id, breaker=breaker, repair=repair, tracer=tracer)
|
||||
# W-AXIS 波1 真相层:逐 turn 全文落 amgen-<gameId>/turns.jsonl(模型文本/工具全参/工具返回/门续修反馈)。
|
||||
# observe-only、best-effort,默认开(CHEAP_TURNS_ENABLED=0 关);接线失败/关 → None,不加入列表(现有行为字节不变)。
|
||||
# RepairMiddleware 在单次 reply 内 mid-reply 注入 name=gate 续修,turns 中间件每见新一轮 ModelCallStart 扫
|
||||
# agent.state.context 收下——故放在 repair 之后无碍(它只读 context 尾部、不拦事件)。
|
||||
mws = [collector, tracer, repair, breaker]
|
||||
# 历史模式下 RepairMiddleware 会 mid-reply 注入 name=gate;v3 没有该事件。turns 只读 context 尾部,
|
||||
# 两种模式都可放最内层,不改变控制流。
|
||||
mws = [collector, tracer]
|
||||
if repair is not None:
|
||||
mws.append(repair)
|
||||
mws.append(breaker)
|
||||
try:
|
||||
import cheap_turns_sink # noqa: PLC0415 —— 顶层零 agentscope,工厂内惰性建中间件
|
||||
turns_mw = cheap_turns_sink.build_turns_middleware(game_id, trace_id=game_id)
|
||||
|
||||
@ -4,7 +4,8 @@ worker_service.py 收 §6.1 job 后,不再进程内 run_studio,而是驱动 chea
|
||||
注册 OpenAI 兼容凭据 → 建 agent(system prompt 的 ⟦G⟧=后端 gameId)/session → scaffold(后端 gameId,起点落
|
||||
amgen-<后端 gameId>)+ 写 session→gameId 映射注册表 sidecar → 设 BYPASS → 发 kick → SSE 等这一次(内部续修多轮)
|
||||
回合真结束 → 读九门 verdict + Service 收口采集 sidecar 组 run-summary → reply 外非阻塞丰富度评分。续修/门判/软预算
|
||||
全在 Service 端 middleware(阶段一① 复用),消费方只发一次 POST。**C2**:产物目录/评门/回调 trace 全用后端 gameId,
|
||||
v3 下旧 gameplay resume 已摘,消费方首回合后跑 v3;仅 verified reject 可向同一 session 再发一次修复 POST。**C2**:
|
||||
产物目录/评门/回调 trace 全用后端 gameId,
|
||||
Service 两工厂经会话注册表把框架分配的 session_id 解析回后端 gameId(工厂拿不到后端 gameId,故靠 sidecar 桥接)。
|
||||
|
||||
【惰性 import 红线】顶层零重依赖;httpx/agentscope 牵出的件在函数体内 import。
|
||||
@ -14,8 +15,99 @@ Service 两工厂经会话注册表把框架分配的 session_id 解析回后端
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import hashlib
|
||||
import json
|
||||
import os
|
||||
import shutil
|
||||
import stat
|
||||
import tempfile
|
||||
import time
|
||||
from pathlib import PurePosixPath
|
||||
|
||||
|
||||
_REFERENCE_RECEIPTS_FILE = ".reference-receipts.json"
|
||||
_SESSION_CFG_WRITE_MAX_BYTES = 1 * 1024 * 1024
|
||||
|
||||
|
||||
def _reference_snapshot_dir(session_cfg_path):
|
||||
"""从服务端固定 sidecar 路径派生 session 独立快照目录,不接受 job 自报路径。"""
|
||||
return session_cfg_path.with_name(f"{session_cfg_path.stem}.reference-assets")
|
||||
|
||||
|
||||
def _validated_snapshot_path(value: str) -> str:
|
||||
"""防御性复核验证器返回路径,确保物化永远留在临时 snapshot 根内。"""
|
||||
if not isinstance(value, str) or not value or "\\" in value or "\x00" in value:
|
||||
raise ValueError("reference snapshot 路径非法")
|
||||
path = PurePosixPath(value)
|
||||
if path.is_absolute() or any(part in ("", ".", "..") for part in path.parts):
|
||||
raise ValueError("reference snapshot 路径越界")
|
||||
return value
|
||||
|
||||
|
||||
def _read_existing_session_cfg_for_write(path):
|
||||
"""写 session-cfg 前读取既有文件,无法确认时统一拒绝覆盖。
|
||||
|
||||
返回 ``None`` 表示文件不存在;返回字典表示已确认是普通 JSON 对象。读取使用固定上限、
|
||||
``O_NOFOLLOW`` 以及打开前后文件身份/元数据复核,避免把损坏文件、特殊文件或竞态中的文件
|
||||
当成可安全修复的旧 sidecar。调用方据此区分「普通无 policy 可兼容覆盖」与「冻结/不可确认必须止损」。
|
||||
"""
|
||||
try:
|
||||
before = os.lstat(path)
|
||||
except FileNotFoundError:
|
||||
return None
|
||||
except OSError as e:
|
||||
raise ValueError("已有 session-cfg 状态不可确认") from e
|
||||
|
||||
# 既有路径必须是稳定的普通文件;symlink 或特殊文件都不能作为覆盖前的安全依据。
|
||||
if stat.S_ISLNK(before.st_mode) or not stat.S_ISREG(before.st_mode):
|
||||
raise ValueError("已有 session-cfg 必须是普通文件且不得为 symlink")
|
||||
if before.st_size > _SESSION_CFG_WRITE_MAX_BYTES:
|
||||
raise ValueError("已有 session-cfg 超过固定读取上限")
|
||||
nofollow = getattr(os, "O_NOFOLLOW", None)
|
||||
if nofollow is None:
|
||||
raise ValueError("当前平台无法确认 session-cfg 非 symlink")
|
||||
|
||||
fd = None
|
||||
try:
|
||||
fd = os.open(path, os.O_RDONLY | nofollow)
|
||||
opened = os.fstat(fd)
|
||||
if (not stat.S_ISREG(opened.st_mode)
|
||||
or (opened.st_dev, opened.st_ino) != (before.st_dev, before.st_ino)):
|
||||
raise ValueError("已有 session-cfg 读取期间发生替换")
|
||||
|
||||
chunks = []
|
||||
total = 0
|
||||
while True:
|
||||
chunk = os.read(
|
||||
fd, min(64 * 1024, _SESSION_CFG_WRITE_MAX_BYTES + 1 - total))
|
||||
if not chunk:
|
||||
break
|
||||
total += len(chunk)
|
||||
if total > _SESSION_CFG_WRITE_MAX_BYTES:
|
||||
raise ValueError("已有 session-cfg 超过固定读取上限")
|
||||
chunks.append(chunk)
|
||||
|
||||
after = os.fstat(fd)
|
||||
if any(getattr(opened, field) != getattr(after, field)
|
||||
for field in ("st_dev", "st_ino", "st_size", "st_mtime_ns", "st_ctime_ns")):
|
||||
raise ValueError("已有 session-cfg 读取期间发生漂移")
|
||||
raw = b"".join(chunks)
|
||||
except OSError as e:
|
||||
raise ValueError(f"已有 session-cfg 读取边界拒绝:{type(e).__name__}") from e
|
||||
finally:
|
||||
if fd is not None:
|
||||
try:
|
||||
os.close(fd)
|
||||
except OSError:
|
||||
pass
|
||||
|
||||
try:
|
||||
cfg = json.loads(raw.decode("utf-8"))
|
||||
except Exception as e: # noqa: BLE001 —— 任何无法确认的 JSON 都不得被覆盖修复
|
||||
raise ValueError("已有 session-cfg JSON 不可确认") from e
|
||||
if not isinstance(cfg, dict):
|
||||
raise ValueError("已有 session-cfg 顶层必须是对象")
|
||||
return cfg
|
||||
|
||||
|
||||
def _resolve_base() -> str:
|
||||
@ -69,7 +161,8 @@ def _cheap_credential_payload(user_token: str | None = None) -> dict:
|
||||
|
||||
def _write_session_cfg(session_id: str, *, external_game_id: str,
|
||||
write_whitelist=None, scaffold_template=None,
|
||||
traceparent=None, tracestate=None) -> None:
|
||||
traceparent=None, tracestate=None,
|
||||
reference_assets=None) -> bool:
|
||||
"""写本 session 的会话注册表 sidecar(C2:Service 两工厂读它把 session_id 解析回后端 gameId + 取 write_whitelist)。
|
||||
|
||||
按 session_id 键写 game-runtime/games/_cheap-sessions/<session_id>.json(worker↔Service 同机共享 FS);
|
||||
@ -77,12 +170,16 @@ def _write_session_cfg(session_id: str, *, external_game_id: str,
|
||||
restricted = 是否受限写(create 路 False + write_whitelist=None;reskin/modify 路 True + 白名单;I1 fail-closed
|
||||
依赖 restricted 标记)。traceparent/tracestate = 面四断点③ 桥:worker 这跳的 W3C context,让工厂据它把
|
||||
cheap-service 生成 span 挂到入站 trace 下(reply 在后台任务跑、读不到出站 header 的实时 context,故走 sidecar 桥)。
|
||||
best-effort:写失败只告警(Service 侧读缺失 → 回落 session_id、生成响亮失败)。
|
||||
无策略路径写失败仍按旧行为只告警;显式策略调用方必须检查 False 并在 /chat 前终止。
|
||||
sidecar 始终通过同目录临时文件 + os.replace 原子发布;显式策略额外先原子发布独立 snapshot 目录。
|
||||
"""
|
||||
import json # noqa: PLC0415
|
||||
|
||||
import cheap_run # noqa: PLC0415
|
||||
import reference_asset_gate # noqa: PLC0415
|
||||
|
||||
snapshot_dir = None
|
||||
snapshot_published = False
|
||||
temp_snapshot = None
|
||||
temp_cfg = None
|
||||
try:
|
||||
# I1 fail-closed:restricted 由「是否给定白名单(is not None)」判,不用 bool()——空集白名单 bool 为 False 会误成
|
||||
# create 的不收窄放开全写;is not None 让空集 → restricted True → T2 读侧收窄成空集禁写(受限会话缺白名单宁禁勿放)。
|
||||
@ -90,6 +187,20 @@ def _write_session_cfg(session_id: str, *, external_game_id: str,
|
||||
wl = sorted(write_whitelist) if write_whitelist else None
|
||||
p = cheap_run.session_cfg_path(session_id)
|
||||
p.parent.mkdir(parents=True, exist_ok=True)
|
||||
existing_cfg = _read_existing_session_cfg_for_write(p)
|
||||
if existing_cfg is not None and "reference_asset_policy" in existing_cfg:
|
||||
# 只要既有 cfg 留下过冻结声明,就算 snapshot 消失也不得用默认写入抹掉证据。
|
||||
raise FileExistsError("已有 session-cfg 含 reference asset policy")
|
||||
snapshot_dir = _reference_snapshot_dir(p)
|
||||
try:
|
||||
os.lstat(snapshot_dir)
|
||||
except FileNotFoundError:
|
||||
pass
|
||||
except OSError as e:
|
||||
raise ValueError("session reference snapshot 状态不可读") from e
|
||||
else:
|
||||
# 同一 session 一旦有冻结 snapshot,默认写入也不得覆盖其 policy 声明。
|
||||
raise FileExistsError("session reference snapshot 已存在")
|
||||
cfg = {
|
||||
"external_game_id": str(external_game_id),
|
||||
"write_whitelist": wl,
|
||||
@ -101,10 +212,83 @@ def _write_session_cfg(session_id: str, *, external_game_id: str,
|
||||
cfg["traceparent"] = traceparent
|
||||
if tracestate:
|
||||
cfg["tracestate"] = tracestate
|
||||
p.write_text(json.dumps(cfg, ensure_ascii=False), encoding="utf-8")
|
||||
|
||||
if reference_assets is not None:
|
||||
if not isinstance(reference_assets, reference_asset_gate.VerifiedReferenceAssets):
|
||||
raise TypeError("reference_assets 必须是 VerifiedReferenceAssets")
|
||||
files = dict(reference_assets.reference_files)
|
||||
total_bytes = sum(len(content) for content in files.values())
|
||||
if total_bytes > reference_asset_gate.MAX_TOTAL_BYTES:
|
||||
raise ValueError("reference snapshot 超过 128 MiB 硬帽")
|
||||
if reference_asset_gate.snapshot_hash(files) != reference_assets.snapshot_hash:
|
||||
raise ValueError("reference snapshot hash 与验证结果不一致")
|
||||
|
||||
temp_snapshot = tempfile.mkdtemp(
|
||||
prefix=f".{p.stem}.reference-assets.", dir=str(p.parent))
|
||||
temp_snapshot_path = p.parent / os.path.basename(temp_snapshot)
|
||||
file_index = []
|
||||
for logical_path in sorted(files, key=lambda item: item.encode("utf-8")):
|
||||
safe_path = _validated_snapshot_path(logical_path)
|
||||
content = files[logical_path]
|
||||
if not isinstance(content, bytes):
|
||||
raise TypeError("reference snapshot 文件必须是 bytes")
|
||||
target = temp_snapshot_path.joinpath(*PurePosixPath(safe_path).parts)
|
||||
target.parent.mkdir(parents=True, exist_ok=True)
|
||||
with target.open("xb") as handle:
|
||||
handle.write(content)
|
||||
handle.flush()
|
||||
os.fsync(handle.fileno())
|
||||
file_index.append({
|
||||
"path": safe_path,
|
||||
"size": len(content),
|
||||
"sha256": hashlib.sha256(content).hexdigest(),
|
||||
})
|
||||
|
||||
receipt_payload = {
|
||||
"receipts": reference_asset_gate.to_json_value(reference_assets.receipts),
|
||||
}
|
||||
receipt_bytes = reference_asset_gate.canonical_json_bytes(receipt_payload)
|
||||
receipt_hash = hashlib.sha256(receipt_bytes).hexdigest()
|
||||
receipt_path = temp_snapshot_path / _REFERENCE_RECEIPTS_FILE
|
||||
with receipt_path.open("xb") as handle:
|
||||
handle.write(receipt_bytes)
|
||||
handle.flush()
|
||||
os.fsync(handle.fileno())
|
||||
|
||||
cfg["reference_asset_policy"] = {
|
||||
"policy_id": "survivor-gold-v1",
|
||||
"mode": "frozen_preflight",
|
||||
"snapshot_hash": reference_assets.snapshot_hash,
|
||||
"receipt_hash": receipt_hash,
|
||||
"roots": reference_asset_gate.to_json_value(reference_assets.reference_roots),
|
||||
"files": file_index,
|
||||
}
|
||||
os.replace(temp_snapshot_path, snapshot_dir)
|
||||
snapshot_published = True
|
||||
temp_snapshot = None
|
||||
|
||||
cfg_bytes = json.dumps(cfg, ensure_ascii=False, separators=(",", ":")).encode("utf-8")
|
||||
fd, temp_cfg = tempfile.mkstemp(prefix=f".{p.name}.", dir=str(p.parent))
|
||||
with os.fdopen(fd, "wb") as handle:
|
||||
handle.write(cfg_bytes)
|
||||
handle.flush()
|
||||
os.fsync(handle.fileno())
|
||||
os.replace(temp_cfg, p)
|
||||
temp_cfg = None
|
||||
return True
|
||||
except Exception as e: # noqa: BLE001 —— sidecar best-effort,写失败 Service 侧回落 session_id
|
||||
print(f"[cheap-driver] session-cfg 写失败(Service 侧将回落 game_id=session_id、生成响亮失败):"
|
||||
f"{type(e).__name__}: {e}", flush=True)
|
||||
if temp_cfg:
|
||||
try:
|
||||
os.unlink(temp_cfg)
|
||||
except OSError:
|
||||
pass
|
||||
if temp_snapshot:
|
||||
shutil.rmtree(temp_snapshot, ignore_errors=True)
|
||||
if snapshot_published and snapshot_dir is not None:
|
||||
shutil.rmtree(snapshot_dir, ignore_errors=True)
|
||||
return False
|
||||
|
||||
|
||||
def _read_last_cheap_verdict(game_id: str):
|
||||
@ -239,8 +423,34 @@ def _failed_summary(game_id: str, reason: str) -> dict:
|
||||
}
|
||||
|
||||
|
||||
async def _run_v3_floor_gates(game_id: str) -> dict:
|
||||
"""Service v3 回合结束后只跑一次机械收口;不在九门结果上做 gameplay resume。"""
|
||||
import asyncio # noqa: PLC0415
|
||||
|
||||
import cheap_gates # noqa: PLC0415
|
||||
import cheap_verify # noqa: PLC0415
|
||||
|
||||
port, cdp_port = cheap_verify._derive_playtest_ports(f"{game_id}:service-floor")
|
||||
try:
|
||||
return await asyncio.to_thread(cheap_gates.run_cheap_gates, game_id, port, cdp_port)
|
||||
except Exception as e: # noqa: BLE001 —— v3 会把缺失四门证据诚实判 tester_error
|
||||
print(f"[cheap-driver] game={game_id} v3 机械门异常:{type(e).__name__}: {e}", flush=True)
|
||||
return {}
|
||||
|
||||
|
||||
def _merge_service_turns(first: dict, second: dict) -> dict:
|
||||
"""合并两次 Service writer 回合成本;第二回合就是 v3 唯一修复。"""
|
||||
out = dict(second or {})
|
||||
out["costRmb"] = round(float((first or {}).get("costRmb") or 0.0)
|
||||
+ float((second or {}).get("costRmb") or 0.0), 4)
|
||||
out["repairs"] = 1
|
||||
out["budgetSoftTripped"] = bool((first or {}).get("budgetSoftTripped")
|
||||
or (second or {}).get("budgetSoftTripped"))
|
||||
return out
|
||||
|
||||
|
||||
async def drive_cheap_generation(job: dict, *, base_url: str | None = None, user_id: str = "cheap"):
|
||||
"""驱动 cheap Service /chat 跑一局便宜档生成(单 POST + 洋葱内续修),返回 (run-summary, game_dir)。
|
||||
"""驱动 cheap Service /chat 跑一局便宜档生成;v3 最多首轮 + 一次 verified-reject 修复。返回摘要与目录。
|
||||
|
||||
与旧 worker_service._default_run_fn 同契约(process_job 零改动消费)。create 路默认:通用 scaffold + 通用系统提示 +
|
||||
write_whitelist=None(对齐现 _default_run_fn 的 run_studio(game_id, brief) 无 scaffold/whitelist)。
|
||||
@ -262,6 +472,7 @@ async def drive_cheap_generation(job: dict, *, base_url: str | None = None, user
|
||||
import cheap_otlp_sink # noqa: PLC0415 —— 面四三跳传播:取当前 context 的 W3C carrier(顶层零重依赖)
|
||||
import cheap_run # noqa: PLC0415
|
||||
import cheap_verify # noqa: PLC0415
|
||||
import cheap_studio # noqa: PLC0415
|
||||
from cheap_roles import build_system_prompt, SCAFFOLD_DESC_BY_TEMPLATE # noqa: PLC0415
|
||||
from service.control_plane import _wait_for_turn_end # noqa: PLC0415 —— 复用 tier2 SSE 等回合
|
||||
|
||||
@ -270,6 +481,33 @@ async def drive_cheap_generation(job: dict, *, base_url: str | None = None, user
|
||||
game_id = str(game_id) if game_id is not None else None
|
||||
brief = job.get("brief") or ""
|
||||
t0 = time.time()
|
||||
acceptance_mode = cheap_studio.acceptance_v3_mode()
|
||||
v3_enabled = acceptance_mode in ("v3", "v3_shadow")
|
||||
reference_asset_policy_id = job.get("referenceAssetPolicyId")
|
||||
frozen_reference_assets = None
|
||||
frozen_reference_constraint_block = None
|
||||
reference_asset_generation_receipts = None
|
||||
if reference_asset_policy_id is not None:
|
||||
try:
|
||||
# 只把服务端 job 的 policyId 作为选择信号;路径、release 和 hash 全由统一 helper 固定。
|
||||
frozen_reference_assets = cheap_verify.preflight_reference_asset_policy(
|
||||
reference_asset_policy_id, acceptance_mode)
|
||||
frozen_reference_constraint_block = cheap_verify.build_frozen_reference_asset_constraint_block(
|
||||
frozen_reference_assets)
|
||||
import reference_asset_gate # noqa: PLC0415
|
||||
reference_asset_generation_receipts = reference_asset_gate.to_json_value(
|
||||
frozen_reference_assets.receipts)
|
||||
except Exception as exc: # noqa: BLE001 可信预检失败不得创建 Writer 请求
|
||||
reason = f"reference asset frozen preflight 失败:{type(exc).__name__}: {exc}"
|
||||
return _failed_summary(game_id, reason), cheap_run.game_dir(game_id)
|
||||
task_binding_hash = None
|
||||
if v3_enabled:
|
||||
try:
|
||||
task_binding_hash = cheap_verify.task_binding_hash_v3(job.get("traceId") or job.get("job_id"))
|
||||
except ValueError as exc:
|
||||
failed = cheap_studio.apply_v3_entry_failure(
|
||||
_failed_summary(game_id, str(exc)), str(exc), mode=acceptance_mode)
|
||||
return failed, cheap_run.game_dir(game_id)
|
||||
|
||||
# WU2 §3.7 F:从 §6.1 job 取 per-user token(后端 dispatchGeneric 对真实 member 装、系统/编排旁路留空);
|
||||
# 空串归一为 None(缺失 → _resolve_key 回落全局 env key)。整条 job 由 worker_service 透传至此,故直接从 job 取。
|
||||
@ -371,18 +609,77 @@ async def drive_cheap_generation(job: dict, *, base_url: str | None = None, user
|
||||
# ④c evidence 清理清单化(W-AXIS 波1 F0-c,取代旧红线③只清 verdict.json 一个文件):清 _wg1-gen/<id>/evidence/
|
||||
# 残留(verdict/续修反馈/日志/截图全列)+ 兜底清 amgen-<id>/evidence/,封「读回上一 run 反馈/绿 verdict → 误归因/假绿」。
|
||||
await asyncio.to_thread(cheap_run.clean_stale_evidence, game_id)
|
||||
acceptance_identity = None
|
||||
if v3_enabled:
|
||||
try:
|
||||
acceptance_identity = cheap_verify.build_acceptance_v3_identity(
|
||||
game_id, brief, genre=str(genre or ""),
|
||||
template_route=str(scaffold_template or ""), repair_ordinal=0,
|
||||
task_binding_hash=task_binding_hash,
|
||||
reference_asset_record_ids=(
|
||||
[receipt["recordId"] for receipt in reference_asset_generation_receipts]
|
||||
if reference_asset_generation_receipts is not None else None),
|
||||
consumer_ref=(reference_asset_gate.POLICY_CONSUMER_REF
|
||||
if reference_asset_generation_receipts is not None else None))
|
||||
except Exception as exc: # noqa: BLE001 —— 无可信 route/profile 时禁止向 Writer 发首条消息
|
||||
failed = _failed_summary(
|
||||
game_id, f"Writer 前无法冻结 v3 acceptance identity:{type(exc).__name__}: {exc}")
|
||||
failed = cheap_studio.apply_v3_entry_failure(
|
||||
failed, failed.get("fail") or "acceptance identity 非法", mode=acceptance_mode)
|
||||
if archived_to:
|
||||
failed["archivedPriorRunTo"] = archived_to
|
||||
return failed, cheap_run.game_dir(game_id)
|
||||
# ⑤ 写会话注册表 sidecar(C2:Service 两工厂据 session_id 读它解析回后端 gameId、绑六工具/评门/collector;
|
||||
# create 路 write_whitelist=None → restricted=False)。external_game_id 恒写、正常路工厂必读到。
|
||||
_write_session_cfg(session_id, external_game_id=game_id,
|
||||
write_whitelist=write_whitelist, scaffold_template=scaffold_template,
|
||||
traceparent=otel_carrier.get("traceparent"),
|
||||
tracestate=otel_carrier.get("tracestate"))
|
||||
sidecar_ok = _write_session_cfg(
|
||||
session_id, external_game_id=game_id,
|
||||
write_whitelist=write_whitelist, scaffold_template=scaffold_template,
|
||||
traceparent=otel_carrier.get("traceparent"),
|
||||
tracestate=otel_carrier.get("tracestate"),
|
||||
reference_assets=frozen_reference_assets,
|
||||
)
|
||||
if not sidecar_ok:
|
||||
if frozen_reference_assets is not None:
|
||||
failed = _failed_summary(game_id, "reference asset session snapshot 原子发布失败")
|
||||
if archived_to:
|
||||
failed["archivedPriorRunTo"] = archived_to
|
||||
return failed, cheap_run.game_dir(game_id)
|
||||
# 旧默认 sidecar 写失败仍保持 best-effort;但发现已有冻结 snapshot 时必须停在 /chat 前。
|
||||
try:
|
||||
os.lstat(_reference_snapshot_dir(cheap_run.session_cfg_path(session_id)))
|
||||
except FileNotFoundError:
|
||||
pass
|
||||
except OSError:
|
||||
failed = _failed_summary(game_id, "已有 reference asset session snapshot 状态不可确认")
|
||||
if archived_to:
|
||||
failed["archivedPriorRunTo"] = archived_to
|
||||
return failed, cheap_run.game_dir(game_id)
|
||||
else:
|
||||
failed = _failed_summary(game_id, "默认 sidecar 不得覆盖已有 reference asset session snapshot")
|
||||
if archived_to:
|
||||
failed["archivedPriorRunTo"] = archived_to
|
||||
return failed, cheap_run.game_dir(game_id)
|
||||
try:
|
||||
existing_cfg = _read_existing_session_cfg_for_write(
|
||||
cheap_run.session_cfg_path(session_id))
|
||||
except ValueError:
|
||||
failed = _failed_summary(game_id, "已有 reference asset session-cfg 状态不可确认")
|
||||
if archived_to:
|
||||
failed["archivedPriorRunTo"] = archived_to
|
||||
return failed, cheap_run.game_dir(game_id)
|
||||
if existing_cfg is not None and "reference_asset_policy" in existing_cfg:
|
||||
failed = _failed_summary(game_id, "默认 sidecar 不得覆盖已有 reference asset policy cfg")
|
||||
if archived_to:
|
||||
failed["archivedPriorRunTo"] = archived_to
|
||||
return failed, cheap_run.game_dir(game_id)
|
||||
# ⑥ 设 BYPASS 权限(六工具默认 ASK,服务态无人确认,不设首个工具调用即卡死;PATCH 需 query agent_id,同 tier2)。
|
||||
await http.patch(f"{base_url}/sessions/{session_id}", json={"permission_mode": "bypass"},
|
||||
headers=headers, params={"agent_id": agent_id}, timeout=30.0)
|
||||
# ⑦ 发 kick(引导 read skill → 写 game-logic.js → check/build → finish;文本对齐旧 run_studio 的 create kick)。
|
||||
kick = (f"请按这个 brief 造一款游戏:「{brief}」。先 read_file 读手册(.agents/skills/littlejs-game-dev.md)"
|
||||
"和你的起点 game-logic.js 再动手;核心玩法实现完、check 与 build 都绿了就立即 finish。")
|
||||
"和你的起点 game-logic.js 再动手;核心玩法实现完、check 与 build 都绿了就立即 finish。"
|
||||
+ (f"\n\n{frozen_reference_constraint_block}"
|
||||
if frozen_reference_constraint_block else ""))
|
||||
await http.post(f"{base_url}/chat/", json={
|
||||
"agent_id": agent_id, "session_id": session_id,
|
||||
"input": {"name": "user", "role": "user", "content": [{"type": "text", "text": kick}]},
|
||||
@ -402,13 +699,73 @@ async def drive_cheap_generation(job: dict, *, base_url: str | None = None, user
|
||||
|
||||
# ⑨ 读九门 verdict + Service 收口采集 sidecar → 组 run-summary(供 result_out;C2:全用后端 gameId)。
|
||||
# _read_service_run_summary 带有界轮询(C1:关闭 collector flush 与 driver 读 REPLY_END 的竞态)。
|
||||
verdict = _read_last_cheap_verdict(game_id)
|
||||
# SSE 未真结束时 writer 可能仍在改盘,禁止并发起机械门;交 v3 以缺失证据诚实判 tester_error。
|
||||
verdict = (await _run_v3_floor_gates(game_id) if v3_enabled and turn.get("ended")
|
||||
else ({} if v3_enabled else _read_last_cheap_verdict(game_id)))
|
||||
svc = _read_service_run_summary(game_id)
|
||||
summary = _build_summary(game_id, brief, verdict, svc, turn, t0)
|
||||
# F0-b:上一 run 归档指针随 run-summary 透传(批次账每 run 记一笔;首跑/关/失败为 None,不带该键)。
|
||||
if archived_to:
|
||||
summary["archivedPriorRunTo"] = archived_to
|
||||
|
||||
if v3_enabled and genre not in cheap_verify._V3_GENRES:
|
||||
summary = cheap_studio.apply_v3_entry_failure(
|
||||
summary, f"v3 无法从 brief/template 元数据确定 canonical genre:{genre or '-'}",
|
||||
mode=acceptance_mode)
|
||||
print(f"[cheap-driver] game={game_id} v3 genre 缺失,冻结发布且不回落旧验收。", flush=True)
|
||||
elif v3_enabled:
|
||||
acceptance = await cheap_verify.run_acceptance_v3(cheap_studio.build_acceptance_v3_request(
|
||||
game_id, brief, verdict, acceptance_identity=acceptance_identity,
|
||||
idempotency_key=f"{job.get('traceId') or game_id}:service-v3",
|
||||
writer_cost_rmb=float(svc.get("costRmb") or 0.0),
|
||||
reference_asset_policy_id=reference_asset_policy_id,
|
||||
reference_asset_generation_receipts=reference_asset_generation_receipts))
|
||||
decision = acceptance.get("decision") or {}
|
||||
acceptance_first_pass = None
|
||||
print(f"[cheap-driver] game={game_id} 验收 v3 mode={acceptance_mode} outcome={decision.get('outcome')} "
|
||||
f"accepted={decision.get('accepted')} publishFrozen={decision.get('publishFrozen')} "
|
||||
f"repairEligible={decision.get('repairEligible')}", flush=True)
|
||||
if cheap_verify.is_v3_repair_authorized(acceptance):
|
||||
# Service v3 不装 RepairMiddleware;只有 final verified reject 才额外 POST 同一会话一次。
|
||||
acceptance_first_pass = acceptance
|
||||
acceptance_identity = cheap_verify.build_acceptance_v3_identity(
|
||||
game_id, brief, genre=acceptance_identity["genre"],
|
||||
template_route=acceptance_identity["templateRoute"],
|
||||
source_artifact_hash=acceptance["artifactHash"],
|
||||
parent_acceptance_request_hash=acceptance_identity["acceptanceRequestHash"],
|
||||
repair_ordinal=1, proof_profile_id=acceptance_identity["proofProfileId"],
|
||||
proof_registry_version=acceptance_identity["proofRegistryVersion"],
|
||||
task_binding_hash=acceptance_identity["taskBindingHash"],
|
||||
design_ref=acceptance_identity.get("designRef"),
|
||||
reference_asset_record_ids=acceptance_identity.get("referenceAssetRecordIds"),
|
||||
consumer_ref=acceptance_identity.get("consumerRef"),
|
||||
)
|
||||
feedback = cheap_studio.build_v3_repair_prompt(decision.get("repairFeedback"))
|
||||
async with httpx.AsyncClient() as http:
|
||||
await http.post(f"{base_url}/chat/", json={
|
||||
"agent_id": agent_id, "session_id": session_id,
|
||||
"input": {"name": "user", "role": "user", "content": [{"type": "text", "text": feedback}]},
|
||||
}, headers=headers, timeout=30.0)
|
||||
repair_turn = await _wait_for_turn_end(base_url, agent_id, session_id, user_id=user_id,
|
||||
timeout_s=sse_timeout, idle_timeout_s=sse_idle_s)
|
||||
# 修复回合若未真结束,同样禁止在 writer 改盘中并发验收,也不把首回合 sidecar 误读成第二回合成本。
|
||||
repaired_verdict = await _run_v3_floor_gates(game_id) if repair_turn.get("ended") else {}
|
||||
repaired_svc = _read_service_run_summary(game_id) if repair_turn.get("ended") else {}
|
||||
svc = _merge_service_turns(svc, repaired_svc)
|
||||
summary = _build_summary(game_id, brief, repaired_verdict, svc, repair_turn, t0)
|
||||
if archived_to:
|
||||
summary["archivedPriorRunTo"] = archived_to
|
||||
acceptance = await cheap_verify.run_acceptance_v3(cheap_studio.build_acceptance_v3_request(
|
||||
game_id, brief, repaired_verdict, acceptance_identity=acceptance_identity,
|
||||
idempotency_key=f"{job.get('traceId') or game_id}:service-v3", repair_count=1,
|
||||
parent_run_id=acceptance.get("runId"),
|
||||
# 首轮成本由 sealed parent decision 读取;请求只提交修复回合 writer 的新增成本。
|
||||
writer_cost_rmb=float(repaired_svc.get("costRmb") or 0.0),
|
||||
reference_asset_policy_id=reference_asset_policy_id,
|
||||
reference_asset_generation_receipts=reference_asset_generation_receipts))
|
||||
verdict = repaired_verdict
|
||||
summary = cheap_studio.apply_acceptance_v3(summary, acceptance, first_pass=acceptance_first_pass)
|
||||
|
||||
# ⑩ reply 外收口:非阻塞丰富度 LLM 评分(additive;不进 verdict、不改 ok;token 不污染生成成本台账)。
|
||||
# genre 透传(T4):品类路由命中时同一次评分 additive 追加品类扩展条目(与 run_studio 的 genre 参数同义)。
|
||||
try:
|
||||
@ -424,14 +781,15 @@ async def drive_cheap_generation(job: dict, *, base_url: str | None = None, user
|
||||
# verdict 传入做 floor 投影;端口缺省按 game_id 派生(driver 收口可能并发,测试员起服避撞)。
|
||||
# run_acceptance 内建 fail-closed 与顶层兜底、绝不抛,additive 写 floor/playtest/judge/acceptanceVersion,
|
||||
# 更新 ok/accepted——result_out 据 accepted 落 status。
|
||||
summary = await cheap_verify.run_acceptance(summary, game_id=game_id, brief=brief, verdict=verdict)
|
||||
_js = summary.get("judge") or {}
|
||||
_pt = summary.get("playtest") or {}
|
||||
print(f"[cheap-driver] game={game_id} 验收 v2 mode={summary.get('acceptanceVersion')} "
|
||||
f"floor={(summary.get('floor') or {}).get('pass')} playtest.accepted={_pt.get('accepted')} "
|
||||
f"rolls={_pt.get('rollCount')} degraded={_pt.get('degraded')} costRmb={_pt.get('costRmb')} "
|
||||
f"→ accepted={summary.get('accepted')} ok={summary.get('ok')}", flush=True)
|
||||
if not v3_enabled:
|
||||
summary = await cheap_verify.run_acceptance(summary, game_id=game_id, brief=brief, verdict=verdict)
|
||||
_js = summary.get("judge") or {}
|
||||
_pt = summary.get("playtest") or {}
|
||||
print(f"[cheap-driver] game={game_id} 历史验收 mode={summary.get('acceptanceVersion')} "
|
||||
f"floor={(summary.get('floor') or {}).get('pass')} playtest.accepted={_pt.get('accepted')} "
|
||||
f"rolls={_pt.get('rollCount')} degraded={_pt.get('degraded')} costRmb={_pt.get('costRmb')} "
|
||||
f"→ accepted={summary.get('accepted')} ok={summary.get('ok')}", flush=True)
|
||||
|
||||
print(f"[cheap-driver] game={game_id} 单 POST 结束: ok={summary['ok']} attempts={summary['attempts']} "
|
||||
print(f"[cheap-driver] game={game_id} Service 驱动结束: ok={summary['ok']} attempts={summary['attempts']} "
|
||||
f"costRmb={summary['costRmb']} wallSec={summary['wallSec']}", flush=True)
|
||||
return summary, cheap_run.game_dir(game_id)
|
||||
|
||||
@ -16,12 +16,13 @@ import hashlib
|
||||
import json
|
||||
import sys
|
||||
import time
|
||||
import uuid
|
||||
from pathlib import Path
|
||||
|
||||
# 跨包 import 兜底 + key 注入(_bootstrap 模块级把 tier2/gen-worker 加进 sys.path)。
|
||||
sys.path.insert(0, str(Path(__file__).resolve().parent)) # → cheap-worker/(CLI 直跑兼容)
|
||||
import _bootstrap # noqa: E402,F401
|
||||
from cheap_roles import build_system_prompt # noqa: E402
|
||||
from cheap_roles import build_system_prompt, SCAFFOLD_DESC_BY_TEMPLATE # noqa: E402
|
||||
from cheap_toolkit import CheapSession, build_toolkit # noqa: E402
|
||||
import cheap_budget # noqa: E402 预算两段式同源工厂(CLI 与生产 Service 读同一配置源)
|
||||
import cheap_run # noqa: E402
|
||||
@ -114,7 +115,7 @@ def _verdict_brief(v) -> dict:
|
||||
return {"pass": v.get("pass"), "failedGates": failed}
|
||||
|
||||
|
||||
def _closeout_gates(game_id, *, port, cdp_port) -> dict:
|
||||
def _closeout_gates(game_id, *, port, cdp_port, interaction_profile_id=None) -> dict:
|
||||
"""CLI 收口门流水线 stage → smoke →九门 play(一次跑齐;回喂循环与收口段共用)。
|
||||
|
||||
与 Service 路 cheap_gates.run_cheap_gates 同序,多回传 staged/smoke_ok/driver_type 供 run-summary 组装
|
||||
@ -127,7 +128,10 @@ def _closeout_gates(game_id, *, port, cdp_port) -> dict:
|
||||
if not st["ok"]:
|
||||
_rec(f"stage FAIL: {st['output'][:300]}")
|
||||
return out
|
||||
sm = cheap_run.smoke(game_id, port=port, cdp_port=cdp_port)
|
||||
sm = cheap_run.smoke(
|
||||
game_id, port=port, cdp_port=cdp_port,
|
||||
interaction_profile_id=interaction_profile_id,
|
||||
)
|
||||
out["smoke_ok"] = sm["ok"]
|
||||
_rec(f"smoke {'PASS' if sm['ok'] else 'FAIL'}(抓 state 供 play-spec)")
|
||||
if not sm["ok"]:
|
||||
@ -190,11 +194,148 @@ def build_trace_source(verdict, driver_type, attempts, stage, model) -> dict:
|
||||
}
|
||||
|
||||
|
||||
def acceptance_v3_mode() -> str:
|
||||
"""返回当前验收模式;只有 v3/v3_shadow 才进入可信证据闭环。"""
|
||||
return str(cheap_verify._acceptance_v3_cfg().get("mode") or "v3_shadow")
|
||||
|
||||
|
||||
def _new_local_acceptance_trace_id(game_id: str) -> str:
|
||||
"""给没有后端 traceId 的直接 CLI 调用生成本次唯一任务标识。
|
||||
|
||||
同一 ``run_studio`` 内的唯一 repair 复用已经冻结的 identity,不会再次调用本函数;下一次 CLI
|
||||
即使复用 gameId 和完全相同的产物,也会得到不同 taskBindingHash,不能命中旧任务证据。
|
||||
"""
|
||||
return f"studio-{game_id}-{uuid.uuid4().hex}"
|
||||
|
||||
|
||||
def _resolve_acceptance_task_trace_id(game_id: str, supplied_trace_id, *, identity_supplied: bool) -> str:
|
||||
"""解析当前调用的可信任务标识;外部 identity 没有原始 traceId 时拒绝继续。"""
|
||||
if supplied_trace_id is not None:
|
||||
# 复用 canonical 非空字符串校验,避免编排层和验收层出现两套 traceId 形状规则。
|
||||
cheap_verify.task_binding_hash_v3(supplied_trace_id)
|
||||
return supplied_trace_id
|
||||
if identity_supplied:
|
||||
raise ValueError("调用方提供 acceptanceIdentity 时必须同时提供当前任务 traceId")
|
||||
return _new_local_acceptance_trace_id(str(game_id))
|
||||
|
||||
|
||||
def build_acceptance_v3_request(game_id: str, brief: str, verdict, *, acceptance_identity,
|
||||
idempotency_key=None, repair_count=None, parent_run_id=None,
|
||||
writer_cost_rmb=0.0, reference_asset_policy_id=None,
|
||||
reference_asset_generation_receipts=None) -> dict:
|
||||
"""组 v3 唯一入口请求;只提交本轮 writer 增量,历史成本由验收边界读取封存父决策。"""
|
||||
identity = dict(acceptance_identity or {})
|
||||
ordinal = int(identity.get("repairOrdinal") or 0)
|
||||
if repair_count is not None and int(repair_count) != ordinal:
|
||||
raise ValueError("repair_count 必须与 acceptanceIdentity.repairOrdinal 一致")
|
||||
policy_present = reference_asset_policy_id is not None
|
||||
receipts_present = reference_asset_generation_receipts is not None
|
||||
if policy_present != receipts_present:
|
||||
raise ValueError("referenceAssetPolicyId/referenceAssetGenerationReceipts 必须成对出现")
|
||||
reference_fields = {}
|
||||
if policy_present:
|
||||
if not isinstance(reference_asset_policy_id, str) or not reference_asset_policy_id:
|
||||
raise ValueError("referenceAssetPolicyId 必须是非空字符串")
|
||||
if (not isinstance(reference_asset_generation_receipts, list)
|
||||
or not reference_asset_generation_receipts
|
||||
or any(not isinstance(receipt, dict)
|
||||
for receipt in reference_asset_generation_receipts)):
|
||||
raise ValueError("referenceAssetGenerationReceipts 必须是非空 JSON 对象数组")
|
||||
for index, receipt in enumerate(reference_asset_generation_receipts):
|
||||
errors = cheap_verify._validate_v3_schema_only(
|
||||
cheap_verify._REFERENCE_ASSET_RECEIPT_SCHEMA, receipt)
|
||||
if errors:
|
||||
raise ValueError(
|
||||
f"referenceAssetGenerationReceipts[{index}] 形状非法:" + ";".join(errors[:4]))
|
||||
# canonical 往返得到与 Writer 前冻结值等价的 plain JSON 副本,避免后续调用方原地篡改。
|
||||
import reference_asset_gate # noqa: PLC0415
|
||||
frozen_receipts = json.loads(reference_asset_gate.canonical_json_bytes(
|
||||
reference_asset_generation_receipts).decode("utf-8"))
|
||||
reference_fields = {
|
||||
"referenceAssetPolicyId": reference_asset_policy_id,
|
||||
"referenceAssetGenerationReceipts": frozen_receipts,
|
||||
}
|
||||
return {
|
||||
"gameId": str(game_id),
|
||||
"brief": brief or "",
|
||||
"acceptanceIdentity": identity,
|
||||
"verdict": verdict or {},
|
||||
"acceptanceMode": acceptance_v3_mode(),
|
||||
"evidenceMode": "native",
|
||||
"idempotencyKey": idempotency_key or f"{game_id}:acceptance-v3",
|
||||
"repairCountAcrossParentChain": ordinal,
|
||||
"parentRunId": parent_run_id,
|
||||
"writerCostRmb": cheap_verify._finite_nonnegative_rmb_v3(writer_cost_rmb, "writerCostRmb"),
|
||||
**reference_fields,
|
||||
}
|
||||
|
||||
|
||||
def apply_acceptance_v3(summary: dict, acceptance: dict, *, first_pass: dict | None = None) -> dict:
|
||||
"""把 v3 终态单向投影进 run-summary,并保留修复前 factual decision 供批账审计。"""
|
||||
out = dict(summary or {})
|
||||
acceptance = acceptance if isinstance(acceptance, dict) else {}
|
||||
compatibility = acceptance.get("compatibility") if isinstance(acceptance.get("compatibility"), dict) else {}
|
||||
for key, value in compatibility.items():
|
||||
if key != "trace":
|
||||
out[key] = value
|
||||
trace = dict(out.get("trace") or {})
|
||||
trace.update(compatibility.get("trace") or {})
|
||||
out["trace"] = trace
|
||||
if isinstance(acceptance.get("floor"), dict):
|
||||
out["floor"] = acceptance["floor"]
|
||||
decision = acceptance.get("decision") if isinstance(acceptance.get("decision"), dict) else {}
|
||||
out["acceptanceV3"] = acceptance
|
||||
# result-out 防重放边界只认本次 run-summary 与 sealed v3 的逐字段镜像;旧 accept 不能授权新 job/bundle。
|
||||
out["acceptanceRunId"] = acceptance.get("runId")
|
||||
out["acceptanceRequestHash"] = acceptance.get("acceptanceRequestHash")
|
||||
out["acceptanceTaskBindingHash"] = acceptance.get("taskBindingHash")
|
||||
out["acceptanceArtifactHash"] = acceptance.get("artifactHash")
|
||||
out["acceptanceBriefHash"] = acceptance.get("briefHash")
|
||||
out["publishFrozen"] = bool(decision.get("publishFrozen", True))
|
||||
first = first_pass if isinstance(first_pass, dict) else acceptance
|
||||
first_decision = first.get("decision") if isinstance(first.get("decision"), dict) else {}
|
||||
repair_attempted = isinstance(first_pass, dict)
|
||||
if repair_attempted:
|
||||
# 只在真实发生 repair 时保留首轮完整对象;终态权威仍只有 acceptanceV3。
|
||||
out["acceptanceV3FirstPass"] = first_pass
|
||||
out["firstPassAccepted"] = first_decision.get("outcome") == "accept" and first_decision.get("accepted") is True
|
||||
out["repairAttempted"] = repair_attempted
|
||||
out["acceptedAfterRepair"] = bool(repair_attempted and decision.get("outcome") == "accept"
|
||||
and decision.get("accepted") is True)
|
||||
return out
|
||||
|
||||
|
||||
def apply_v3_entry_failure(summary: dict, reason: str, *, mode: str) -> dict:
|
||||
"""入口缺 canonical genre/完整 payload 时显式冻结;声明 v3 后禁止回落旧 accepted。"""
|
||||
out = dict(summary or {})
|
||||
out.update({"acceptanceVersion": mode, "accepted": False, "ok": False, "publishFrozen": True,
|
||||
"failureReason": reason, "repairAttempted": False, "acceptedAfterRepair": False})
|
||||
out["failureLayer"] = {"layer": "tester_degraded", "reason": reason, "failedGates": []}
|
||||
return out
|
||||
|
||||
|
||||
def build_v3_repair_prompt(feedback: str) -> str:
|
||||
"""把 v3 硬证反馈包装成一次性 writer 指令;禁止顺手重做,并强制重过 check/build/finish。"""
|
||||
return ((feedback or "v3 已验证拒绝,请按硬证修复玩法闭环")
|
||||
+ "\n只修上述硬证指向的问题,不要扩需求;完成后必须重新 check → build → finish。")
|
||||
|
||||
|
||||
def next_v3_writer_cost(writer_cost_now: float, writer_cost_accounted: float) -> float:
|
||||
"""只返回尚未计入验收的 writer 增量;父链历史成本不再由调用方搬运。"""
|
||||
current = cheap_verify._finite_nonnegative_rmb_v3(writer_cost_now, "writerCostRmb.current")
|
||||
accounted = cheap_verify._finite_nonnegative_rmb_v3(
|
||||
writer_cost_accounted, "writerCostRmb.accounted")
|
||||
return max(0.0, current - accounted)
|
||||
|
||||
|
||||
async def run_studio(game_id, brief, *, max_iters=40, max_resumes=6, max_tokens=16000,
|
||||
port=4320, cdp_port=9222, run_gates=True,
|
||||
interaction_profile_id=None,
|
||||
system_prompt=None, initial_kick=None, write_whitelist=None, prepare=None,
|
||||
scaffold_template=None, scaffold_desc=None, verify_richness_enabled=True,
|
||||
genre=None):
|
||||
genre=None, acceptance_parent_run_id=None, acceptance_repair_count=0,
|
||||
acceptance_brief=None, acceptance_identity=None, acceptance_template_route=None,
|
||||
acceptance_task_trace_id=None, reference_asset_policy_id=None):
|
||||
"""跑便宜档一局生成(scaffold → ReAct 写 src/ → done 门 → 收口 stage+smoke+九门)。返回 run-summary dict。
|
||||
|
||||
genre(W-GENRE 件④,缺省 None=原行为):品类键(如 'puzzle')显式透传给丰富度 LLM 评分——同一次
|
||||
@ -206,6 +347,10 @@ async def run_studio(game_id, brief, *, max_iters=40, max_resumes=6, max_tokens=
|
||||
**不**跑内部九门 play —— 由调用方在生成与 play 之间注入金标 spec 再单独 play,把驱动器从对照变量里摘掉
|
||||
(Codex C1:run_studio 内部已 play,对照需把生成与 play 拆开)。
|
||||
|
||||
interaction_profile_id(缺省 None=原行为):由可信 scaffold 固化进 Writer 不可写的 entry-bundle,普通浏览器
|
||||
直接打开时据此装配受保护 producer;同时透传给 smoke runner 作为显式覆盖。标准 Match-3 生成批传
|
||||
`match3.orthogonal-swap-v1`。它不写入 Writer 上下文,也不改变其它品类的默认入口或验收身份。
|
||||
|
||||
A11 M4 模块重生成复用本编排(同一 ReAct + resume + 熔断 + 三层校验收口,不另造),靠四个可选参切到 modify 态,
|
||||
都默认 None=create 原行为(create 路零改动):
|
||||
· system_prompt:None=create 的 build_system_prompt;modify 传 build_modify_system_prompt(只改玩法)。
|
||||
@ -213,16 +358,51 @@ async def run_studio(game_id, brief, *, max_iters=40, max_resumes=6, max_tokens=
|
||||
· write_whitelist:None=create 不收窄;modify 传 {"game-logic.js"} 把写边界收窄到只许写玩法文件。
|
||||
· prepare:None=create 的 scaffold(clone _template);modify 传「已由上游 scaffold+materialize base 源」
|
||||
的 noop(game_dir 已就绪、不要再 scaffold 覆盖掉 base 源)。签名同 scaffold:(game_id)->{ok,output}。
|
||||
· acceptance_parent_run_id / acceptance_repair_count:只给 A11 的 v3 修复复验传 parentRun 血缘。
|
||||
· acceptance_brief:writer 指令与验收题面不同时显式传原 brief,防修复反馈污染 briefHash。
|
||||
· reference_asset_policy_id:可信调用方显式选择冻结参照策略;None 时不触发 gate、不改变 session/kick。
|
||||
"""
|
||||
t0 = time.time()
|
||||
_rec(f"model={_bootstrap.SPIKE_MODEL} id={game_id} brief=「{brief}」max_iters={max_iters} max_resumes={max_resumes}")
|
||||
|
||||
acceptance_mode = acceptance_v3_mode()
|
||||
frozen_reference_assets = None
|
||||
frozen_reference_constraint_block = None
|
||||
reference_asset_generation_receipts = None
|
||||
if reference_asset_policy_id is not None:
|
||||
# 可信消费必须先于 scaffold/Writer;helper 内固定仓根、release 与 hash,调用方不能自报信任材料。
|
||||
frozen_reference_assets = cheap_verify.preflight_reference_asset_policy(
|
||||
reference_asset_policy_id, acceptance_mode)
|
||||
frozen_reference_constraint_block = cheap_verify.build_frozen_reference_asset_constraint_block(
|
||||
frozen_reference_assets)
|
||||
import reference_asset_gate # noqa: PLC0415
|
||||
reference_asset_generation_receipts = reference_asset_gate.to_json_value(
|
||||
frozen_reference_assets.receipts)
|
||||
v3_enabled = run_gates and acceptance_mode in ("v3", "v3_shadow")
|
||||
v3_brief = brief if acceptance_brief is None else acceptance_brief
|
||||
# create 的可信路由必须发生在 scaffold/Writer 前;路由一旦选定就同时决定实际模板与 proof profile。
|
||||
# 无法命中五个可信模板时 fail-closed,不允许先用通用模板生成、验收时再按 genre 猜 profile。
|
||||
if v3_enabled and prepare is None and scaffold_template is None:
|
||||
from cheap_genre_route import route_genre # noqa: PLC0415
|
||||
|
||||
scaffold_template, routed_genre = route_genre(v3_brief)
|
||||
genre = genre or routed_genre
|
||||
if scaffold_template and scaffold_desc is None:
|
||||
scaffold_desc = SCAFFOLD_DESC_BY_TEMPLATE.get(scaffold_template)
|
||||
if v3_enabled and acceptance_identity is None and not (acceptance_template_route or scaffold_template):
|
||||
return apply_v3_entry_failure(
|
||||
{"ok": False, "gameId": game_id, "brief": brief, "finished": False},
|
||||
"Writer 前无法选择可信 templateRoute;v3 禁止使用通用模板或仅凭 genre 反推 profile",
|
||||
mode=acceptance_mode,
|
||||
)
|
||||
|
||||
# W-AXIS 波1 F0-b per-run 归档:create 路 scaffold 会 rmSync 整个 amgen-<id>/、重跑同 gid 抹掉上一 run 的
|
||||
# trace.jsonl/turns.jsonl/证据。故 scaffold 前先把上一 run 整体移进带时间戳归档位(永不覆盖)。modify 路
|
||||
# (prepare 给定)base 源须保留、不归档。best-effort:关/失败退回旧覆盖行为(archived_to=None)。
|
||||
archived_to = cheap_run.archive_prior_run(game_id) if prepare is None else None
|
||||
# create=scaffold clone 模板(扩模板:scaffold_template 传 per-genre 黄金骨架名,如经营=_template-shop);modify=上游已预备的 noop
|
||||
prep = prepare or (lambda gid: cheap_run.scaffold(gid, scaffold_template))
|
||||
prep = prepare or (lambda gid: cheap_run.scaffold(
|
||||
gid, scaffold_template, interaction_profile_id=interaction_profile_id))
|
||||
sc = prep(game_id)
|
||||
if not sc["ok"]:
|
||||
return {"ok": False, "gameId": game_id, "fail": "准备失败:" + sc["output"]}
|
||||
@ -233,7 +413,95 @@ async def run_studio(game_id, brief, *, max_iters=40, max_resumes=6, max_tokens=
|
||||
if prepare is None:
|
||||
cheap_run.clean_stale_evidence(game_id)
|
||||
|
||||
session = CheapSession(game_id=game_id)
|
||||
# v3 身份必须在 Writer 创建前由可信编排层冻结。templateRoute 是 profile 选择轴;genre 只能交叉核对,
|
||||
# 绝不允许在验收阶段仅凭 genre 重新猜 profile。modify 路由调用方通过 acceptance_template_route
|
||||
# 或完整 acceptance_identity 传入源项目可信路由,身份仅存本函数局部变量。
|
||||
trusted_route = acceptance_template_route or scaffold_template
|
||||
trusted_genre = genre or (cheap_verify.GENRE_BY_TEMPLATE.get(trusted_route) if trusted_route else None)
|
||||
v3_identity = None
|
||||
if v3_enabled:
|
||||
try:
|
||||
# 生产 Service 透传后端 traceId;直接 CLI 没有后端任务时,在 Writer 前生成本次唯一标识。
|
||||
# 外部 identity 若缺原始 traceId 无法证明属于当前任务,必须 fail-closed。
|
||||
local_task_trace_id = _resolve_acceptance_task_trace_id(
|
||||
str(game_id), acceptance_task_trace_id,
|
||||
identity_supplied=acceptance_identity is not None)
|
||||
expected_task_binding = cheap_verify.task_binding_hash_v3(local_task_trace_id)
|
||||
if acceptance_identity is not None:
|
||||
supplied = dict(acceptance_identity)
|
||||
if expected_task_binding is not None and supplied.get("taskBindingHash") != expected_task_binding:
|
||||
raise ValueError("acceptanceIdentity.taskBindingHash 与当前任务 traceId 不一致")
|
||||
expected = cheap_verify.build_acceptance_v3_identity(
|
||||
game_id, v3_brief, genre=str(supplied.get("genre") or ""),
|
||||
template_route=str(supplied.get("templateRoute") or ""),
|
||||
source_artifact_hash=supplied.get("sourceArtifactHash"),
|
||||
parent_acceptance_request_hash=supplied.get("parentAcceptanceRequestHash"),
|
||||
repair_ordinal=int(supplied.get("repairOrdinal") or 0),
|
||||
proof_profile_id=supplied.get("proofProfileId"),
|
||||
proof_registry_version=supplied.get("proofRegistryVersion"),
|
||||
task_binding_hash=supplied.get("taskBindingHash"),
|
||||
# W-GOLD-LIVE 检查点 3a:外部身份的参照资产消费声明三字段原样透传重建,
|
||||
# 供下方全等比对与消费对账闸消费;缺省 None/[]/null ≡ 未声明(旧行为不变)。
|
||||
design_ref=supplied.get("designRef"),
|
||||
reference_asset_record_ids=supplied.get("referenceAssetRecordIds"),
|
||||
consumer_ref=supplied.get("consumerRef"),
|
||||
)
|
||||
if supplied != expected:
|
||||
raise ValueError("调用方 acceptanceIdentity 与 canonical registry 不一致")
|
||||
v3_identity = supplied
|
||||
else:
|
||||
if int(acceptance_repair_count or 0) != 0:
|
||||
raise ValueError("修复入口必须显式传入锁定 parent/profile/registry 的 acceptanceIdentity")
|
||||
v3_identity = cheap_verify.build_acceptance_v3_identity(
|
||||
game_id, v3_brief, genre=str(trusted_genre or ""),
|
||||
template_route=str(trusted_route or ""), repair_ordinal=0,
|
||||
task_binding_hash=expected_task_binding,
|
||||
reference_asset_record_ids=(
|
||||
[receipt["recordId"] for receipt in reference_asset_generation_receipts]
|
||||
if reference_asset_generation_receipts is not None else None),
|
||||
consumer_ref=(reference_asset_generation_receipts[0]["consumerRef"]
|
||||
if reference_asset_generation_receipts is not None else None))
|
||||
except Exception as exc: # noqa: BLE001 —— 无可信身份时禁止 Writer 改盘
|
||||
return apply_v3_entry_failure(
|
||||
{"ok": False, "gameId": game_id, "brief": brief, "finished": False},
|
||||
f"Writer 前无法冻结 v3 acceptance identity:{type(exc).__name__}: {exc}",
|
||||
mode=acceptance_mode,
|
||||
)
|
||||
|
||||
if frozen_reference_assets is not None and v3_identity is not None:
|
||||
# 外部 acceptanceIdentity 即使自身 canonical,也必须与本次 /2 冻结回执逐项一致;
|
||||
# 否则会让 Writer 在验收侧最终拒绝之前先消费一套未进入 identity 的参照字节。
|
||||
frozen_record_ids = [receipt["recordId"] for receipt in reference_asset_generation_receipts]
|
||||
frozen_consumers = {receipt["consumerRef"] for receipt in reference_asset_generation_receipts}
|
||||
if (v3_identity.get("referenceAssetRecordIds") != frozen_record_ids
|
||||
or len(frozen_consumers) != 1
|
||||
or v3_identity.get("consumerRef") not in frozen_consumers):
|
||||
return apply_v3_entry_failure(
|
||||
{"ok": False, "gameId": game_id, "brief": brief, "finished": False},
|
||||
"Writer 前冻结 policy 与 acceptance identity 不一致",
|
||||
mode=acceptance_mode,
|
||||
)
|
||||
|
||||
# Registry/1 identity 消费闸仅作旧 acceptance 回放兼容;/2 生产消费已在 scaffold 前完成冻结预检,
|
||||
# 其 consumerRef 与 /1 历史值不同,已有冻结快照时不得再进入 /1 对账。
|
||||
v3_reference_constraint_block = None
|
||||
if v3_enabled and v3_identity is not None and frozen_reference_assets is None:
|
||||
try:
|
||||
reference_gate = cheap_verify.build_v3_reference_asset_generation_constraints(v3_identity)
|
||||
except ValueError as exc: # noqa: BLE001 —— 声明消费而对账失败:fail-closed 拒绝生成入口
|
||||
return apply_v3_entry_failure(
|
||||
{"ok": False, "gameId": game_id, "brief": brief, "finished": False},
|
||||
f"参照资产消费对账失败,生成入口拒绝:{exc}",
|
||||
mode=acceptance_mode,
|
||||
)
|
||||
if reference_gate is not None:
|
||||
v3_reference_constraint_block = reference_gate["constraint_block"]
|
||||
|
||||
session = CheapSession(
|
||||
game_id=game_id,
|
||||
reference_files=(frozen_reference_assets.reference_files if frozen_reference_assets else None),
|
||||
reference_roots=(frozen_reference_assets.reference_roots if frozen_reference_assets else None),
|
||||
)
|
||||
toolkit = build_toolkit(session, write_whitelist=write_whitelist) # modify 收窄到只许写 game-logic.js
|
||||
model = _bootstrap.build_cheap_model(max_tokens=max_tokens)
|
||||
# 熔断:便宜档预算两段式(W-ARCH②,创始人 2026-07-03 裁决)——软停线 ¥10(越线只许收尾类动作,
|
||||
@ -289,14 +557,26 @@ async def run_studio(game_id, brief, *, max_iters=40, max_resumes=6, max_tokens=
|
||||
|
||||
kick = initial_kick or (
|
||||
f"请按这个 brief 造一款游戏:「{brief}」。先 read_file 读手册(.agents/skills/littlejs-game-dev.md)"
|
||||
"和你的起点 game-logic.js 再动手;核心玩法实现完、check 与 build 都绿了就立即 finish。")
|
||||
"和你的起点 game-logic.js 再动手;核心玩法实现完、check 与 build 都绿了就立即 finish。"
|
||||
# 旧 Registry/1 identity 约束与 v2 冻结策略约束均只在各自显式选择后追加;None 路径文本不变。
|
||||
+ (f"\n\n{v3_reference_constraint_block}" if v3_reference_constraint_block else "")
|
||||
+ (f"\n\n{frozen_reference_constraint_block}" if frozen_reference_constraint_block else ""))
|
||||
breaker_tripped = None
|
||||
attempts = 0
|
||||
closeout = None # CLI 回喂:最近一次九门收口结果(回喂循环产出,收口段复用、不重复跑门)
|
||||
acceptance_v3 = None
|
||||
acceptance_v3_first_pass = None
|
||||
v3_entry_error = None
|
||||
v3_repair_count = int((v3_identity or {}).get("repairOrdinal") or acceptance_repair_count or 0)
|
||||
v3_parent_run_id = acceptance_parent_run_id
|
||||
# 已计入父链的 writer 成本基线;每轮验收只加自上次验收后的新增 writer delta。
|
||||
v3_writer_cost_accounted = 0.0
|
||||
# v3 的唯一一次 writer 修复不占旧 resume 配额;旧配额继续只服务 check/build 与畸形工具调用恢复。
|
||||
total_attempts = max_resumes + 1 + (1 if v3_enabled and v3_repair_count == 0 else 0)
|
||||
try:
|
||||
for attempt in range(max_resumes + 1):
|
||||
for attempt in range(total_attempts):
|
||||
attempts = attempt + 1
|
||||
_rec(f"resume attempt {attempts}/{max_resumes + 1} → writer.reply …")
|
||||
_rec(f"resume attempt {attempts}/{total_attempts} → writer.reply …")
|
||||
try:
|
||||
await writer.reply(_user_msg(kick))
|
||||
except Tier2CircuitBreak as _cb:
|
||||
@ -314,19 +594,61 @@ async def run_studio(game_id, brief, *, max_iters=40, max_resumes=6, max_tokens=
|
||||
kick = _MALFORMED_RESUME_FEEDBACK
|
||||
continue
|
||||
raise # 其它熔断 / 预算耗尽 / 非漏参 stuck → 终态,交外层 except 记 breaker_tripped
|
||||
# ① 真 finish:CLI 回喂对齐(W-S1 单③,对账发现六④)——收敛判据从「check+build 绿」升为
|
||||
# 「九门绿」:finish 后跑九门收口,门未全绿且轮数/预算有余时带 C6 反馈 resume 续修,
|
||||
# 使 lab 复验口径与生产 RepairMiddleware(finish 点拦截→跑门→judge→回喂)行为同型。
|
||||
# ① 真 finish:先跑机械收口。v3 只听 final decision,九门失败本身不再触发玩法续修;
|
||||
# 显式历史 v1/v2 回放才保留旧 C6 九门反馈 resume。
|
||||
if session.finished is not None:
|
||||
if not run_gates:
|
||||
_rec("finish 已接受(check+build 绿)→ 收敛(generation-only,不跑九门回喂)")
|
||||
break
|
||||
_rec("finish 已接受(check+build 绿)→ 九门收口判(CLI 回喂对齐)")
|
||||
closeout = _closeout_gates(game_id, port=port, cdp_port=cdp_port)
|
||||
j = judge_cheap_verdict(closeout["verdict"], game_id=game_id,
|
||||
staged_dir=cheap_run.wg1_game_dir(game_id))
|
||||
closeout = _closeout_gates(
|
||||
game_id, port=port, cdp_port=cdp_port,
|
||||
interaction_profile_id=interaction_profile_id,
|
||||
)
|
||||
vb = _verdict_brief(closeout["verdict"])
|
||||
_rec(f"九门 play pass={vb['pass']} failedGates={vb['failedGates']}")
|
||||
if v3_enabled:
|
||||
writer_cost_now = float(getattr(breaker, "spent_rmb", 0.0) or 0.0)
|
||||
request_writer_cost = next_v3_writer_cost(
|
||||
writer_cost_now, v3_writer_cost_accounted)
|
||||
acceptance_v3 = await cheap_verify.run_acceptance_v3(build_acceptance_v3_request(
|
||||
game_id, v3_brief, closeout["verdict"], acceptance_identity=v3_identity,
|
||||
idempotency_key=f"{game_id}:studio-v3", repair_count=v3_repair_count,
|
||||
parent_run_id=v3_parent_run_id, writer_cost_rmb=request_writer_cost,
|
||||
reference_asset_policy_id=reference_asset_policy_id,
|
||||
reference_asset_generation_receipts=reference_asset_generation_receipts))
|
||||
decision = acceptance_v3.get("decision") or {}
|
||||
_rec(f"验收 v3 mode={acceptance_mode} outcome={decision.get('outcome')} "
|
||||
f"accepted={decision.get('accepted')} publishFrozen={decision.get('publishFrozen')} "
|
||||
f"repairEligible={decision.get('repairEligible')}")
|
||||
if cheap_verify.is_v3_repair_authorized(acceptance_v3) and v3_repair_count == 0:
|
||||
# 只有 final postguard 产出的 verified reject 才能到这里;清掉旧终态后让同一 writer 修一次。
|
||||
acceptance_v3_first_pass = acceptance_v3
|
||||
# 同一 Writer 继续前先冻结修复身份:profile/registry/templateRoute 原样锁定,
|
||||
# 只把 parent request 与首轮最终 artifact 作为 ordinal=1 血缘写入。
|
||||
v3_identity = cheap_verify.build_acceptance_v3_identity(
|
||||
game_id, v3_brief, genre=v3_identity["genre"],
|
||||
template_route=v3_identity["templateRoute"],
|
||||
source_artifact_hash=acceptance_v3["artifactHash"],
|
||||
parent_acceptance_request_hash=v3_identity["acceptanceRequestHash"],
|
||||
repair_ordinal=1, proof_profile_id=v3_identity["proofProfileId"],
|
||||
proof_registry_version=v3_identity["proofRegistryVersion"],
|
||||
task_binding_hash=v3_identity["taskBindingHash"],
|
||||
design_ref=v3_identity.get("designRef"),
|
||||
reference_asset_record_ids=v3_identity.get("referenceAssetRecordIds"),
|
||||
consumer_ref=v3_identity.get("consumerRef"),
|
||||
)
|
||||
v3_repair_count = v3_identity["repairOrdinal"]
|
||||
v3_parent_run_id = acceptance_v3.get("runId")
|
||||
v3_writer_cost_accounted = writer_cost_now
|
||||
kick = build_v3_repair_prompt(decision.get("repairFeedback"))
|
||||
session.finished = None
|
||||
acceptance_v3 = None # 产物即将变化,旧 artifactHash 的决策不得冒充终态。
|
||||
_rec("v3 verified reject → 同一 writer 仅一次修复;完成后重跑机械门与 v3")
|
||||
continue
|
||||
break
|
||||
j = judge_cheap_verdict(closeout["verdict"], game_id=game_id,
|
||||
staged_dir=cheap_run.wg1_game_dir(game_id))
|
||||
if j.passed:
|
||||
break
|
||||
if attempt >= max_resumes:
|
||||
@ -379,7 +701,10 @@ async def run_studio(game_id, brief, *, max_iters=40, max_resumes=6, max_tokens=
|
||||
driver_type = None # M3b U1:ensure_play_spec 产的 driver 类型(对照路 run_gates=False 不产 → 保持 None)
|
||||
if finished:
|
||||
if run_gates:
|
||||
cg = closeout or _closeout_gates(game_id, port=port, cdp_port=cdp_port) # 防御:正常路循环内已跑
|
||||
cg = closeout or _closeout_gates(
|
||||
game_id, port=port, cdp_port=cdp_port,
|
||||
interaction_profile_id=interaction_profile_id,
|
||||
) # 防御:正常路循环内已跑
|
||||
staged = cg["staged"]
|
||||
smoke_ok = cg["smoke_ok"]
|
||||
driver_type = cg["driver_type"] # M3b U1:driver 类型存进 trace.gatespec.driver(后端 D11 firstPlay 维)
|
||||
@ -389,7 +714,10 @@ async def run_studio(game_id, brief, *, max_iters=40, max_resumes=6, max_tokens=
|
||||
staged = st["ok"]
|
||||
_rec(f"stage {'OK' if staged else 'FAIL: ' + st['output']}")
|
||||
if staged:
|
||||
sm = cheap_run.smoke(game_id, port=port, cdp_port=cdp_port)
|
||||
sm = cheap_run.smoke(
|
||||
game_id, port=port, cdp_port=cdp_port,
|
||||
interaction_profile_id=interaction_profile_id,
|
||||
)
|
||||
smoke_ok = sm["ok"]
|
||||
_rec(f"smoke {'PASS' if smoke_ok else 'FAIL'}(抓 state 供 play-spec)")
|
||||
if not smoke_ok:
|
||||
@ -454,19 +782,36 @@ async def run_studio(game_id, brief, *, max_iters=40, max_resumes=6, max_tokens=
|
||||
"reason": f"richness 接线异常:{type(e).__name__}: {e}"}
|
||||
_rec(f"丰富度评分接线异常(已降级,不影响 run):{type(e).__name__}: {e}")
|
||||
|
||||
# ── W-AXIS-V2 波1:统一验收编排器(floor 四门投影 ∧ 测试 agent 真玩;acceptance.mode 三态)──
|
||||
# ── 统一验收入口:v3 可信证据闭环;仅历史 v1/v2 模式保留旧编排器 ──
|
||||
# 拆着杀:契约无关四门(A/B/C/D 投影)= 预筛权威(取代旧 verdict.pass 九门口径,那随契约退役会坍缩),
|
||||
# 过筛者交测试 agent 视觉引导真玩裁 broken/hollow/off-brief;mode=v2 阻断、shadow 灰度对照、v1 旧口径。
|
||||
# run_acceptance 内建 fail-closed 与顶层兜底、绝不抛,additive 写 summary['floor']/['playtest']/['judge']
|
||||
# /['acceptanceVersion'] 并更新 ['ok']/['accepted']。只在 run_gates(有真 verdict)时接入;对照路零改动。
|
||||
if run_gates and finished:
|
||||
summary = await cheap_verify.run_acceptance(summary, game_id=game_id, brief=brief, verdict=verdict)
|
||||
_js = summary.get("judge") or {}
|
||||
_pt = summary.get("playtest") or {}
|
||||
_rec(f"验收 v2 mode={summary.get('acceptanceVersion')} floor={(summary.get('floor') or {}).get('pass')} "
|
||||
f"playtest.accepted={_pt.get('accepted')} rolls={_pt.get('rollCount')} degraded={_pt.get('degraded')} "
|
||||
f"costRmb={_pt.get('costRmb')} judge.verdict={_js.get('verdict')} "
|
||||
f"→ accepted={summary.get('accepted')} ok={summary.get('ok')}")
|
||||
if v3_enabled:
|
||||
if v3_entry_error:
|
||||
summary = apply_v3_entry_failure(summary, v3_entry_error, mode=acceptance_mode)
|
||||
else:
|
||||
# 正常路径已在 finish 分支完成 v3;防御性补跑只覆盖没有进入该分支的异常控制流。
|
||||
if acceptance_v3 is None:
|
||||
writer_cost_now = float(getattr(breaker, "spent_rmb", 0.0) or 0.0)
|
||||
acceptance_v3 = await cheap_verify.run_acceptance_v3(build_acceptance_v3_request(
|
||||
game_id, v3_brief, verdict, acceptance_identity=v3_identity,
|
||||
idempotency_key=f"{game_id}:studio-v3",
|
||||
repair_count=v3_repair_count, parent_run_id=v3_parent_run_id,
|
||||
writer_cost_rmb=next_v3_writer_cost(
|
||||
writer_cost_now, v3_writer_cost_accounted),
|
||||
reference_asset_policy_id=reference_asset_policy_id,
|
||||
reference_asset_generation_receipts=reference_asset_generation_receipts))
|
||||
summary = apply_acceptance_v3(summary, acceptance_v3, first_pass=acceptance_v3_first_pass)
|
||||
else:
|
||||
summary = await cheap_verify.run_acceptance(summary, game_id=game_id, brief=brief, verdict=verdict)
|
||||
_js = summary.get("judge") or {}
|
||||
_pt = summary.get("playtest") or {}
|
||||
_rec(f"历史验收 mode={summary.get('acceptanceVersion')} floor={(summary.get('floor') or {}).get('pass')} "
|
||||
f"playtest.accepted={_pt.get('accepted')} rolls={_pt.get('rollCount')} degraded={_pt.get('degraded')} "
|
||||
f"costRmb={_pt.get('costRmb')} judge.verdict={_js.get('verdict')} "
|
||||
f"→ accepted={summary.get('accepted')} ok={summary.get('ok')}")
|
||||
|
||||
ev_dir = cheap_run.game_dir(game_id) / "evidence"
|
||||
ev_dir.mkdir(parents=True, exist_ok=True)
|
||||
|
||||
@ -7,16 +7,147 @@ Python 实现、check/build shell-out node。done 门 = finish 工具(工具
|
||||
"""
|
||||
|
||||
import json
|
||||
from dataclasses import dataclass
|
||||
import unicodedata
|
||||
from collections.abc import Mapping
|
||||
from dataclasses import dataclass, field
|
||||
from types import MappingProxyType
|
||||
from typing import Optional
|
||||
|
||||
import cheap_run
|
||||
|
||||
# 单次工具结果注入上限(对 gen.mjs TOOL_RESULT_CAP:够一份 skill/api.d.ts,控 context)。
|
||||
_TOOL_RESULT_CAP = 30000
|
||||
_REFERENCE_READ_BYTES = getattr(cheap_run, "_MAX_READ_BYTES", 200 * 1024)
|
||||
_REFERENCE_DENY_MESSAGE = "ERROR: 受保护参照路径未在只读快照中"
|
||||
|
||||
|
||||
@dataclass
|
||||
def _normalize_repo_path(path: str) -> Optional[str]:
|
||||
"""把仓内逻辑路径规范化为 NFC POSIX 段,拒绝绝对路径和越过仓根的路径。"""
|
||||
if not isinstance(path, str) or not path or "\x00" in path or path.startswith(("/", "\\")):
|
||||
return None
|
||||
path = unicodedata.normalize("NFC", path)
|
||||
parts = []
|
||||
for part in path.split("/"):
|
||||
if part in ("", "."):
|
||||
continue
|
||||
if part == "..":
|
||||
if not parts:
|
||||
return None
|
||||
parts.pop()
|
||||
continue
|
||||
parts.append(part)
|
||||
return "/".join(parts) or None
|
||||
|
||||
|
||||
def _path_requires_rejection(path: str) -> bool:
|
||||
"""判断绝对路径、非法类型或越过仓根的别名,避免它们回落到活目录。"""
|
||||
if not isinstance(path, str) or "\x00" in path or path.startswith(("/", "\\")):
|
||||
return True
|
||||
parts = []
|
||||
for part in path.split("/"):
|
||||
if part in ("", "."):
|
||||
continue
|
||||
if part == "..":
|
||||
if not parts:
|
||||
return True
|
||||
parts.pop()
|
||||
continue
|
||||
parts.append(part)
|
||||
return False
|
||||
|
||||
|
||||
def _path_is_within(path: str, root: str) -> bool:
|
||||
"""按路径段判断包含关系,避免 assets/gold 误匹配 assets/golden。"""
|
||||
return path == root or path.startswith(root + "/")
|
||||
|
||||
|
||||
def _freeze_reference_files(files: Optional[Mapping[str, bytes]]) -> Mapping[str, bytes]:
|
||||
"""复制并冻结验证器给出的 bytes 快照,避免调用方改写 Toolkit 输入。"""
|
||||
if files is None:
|
||||
return MappingProxyType({})
|
||||
if not isinstance(files, Mapping):
|
||||
raise TypeError("reference_files 必须是路径到 bytes 的映射")
|
||||
frozen = {}
|
||||
for raw_path, content in files.items():
|
||||
path = _normalize_repo_path(raw_path)
|
||||
if path is None or not isinstance(content, bytes):
|
||||
raise TypeError("reference_files 必须包含规范化路径和 bytes 内容")
|
||||
frozen[path] = content
|
||||
return MappingProxyType(frozen)
|
||||
|
||||
|
||||
def _freeze_reference_roots(roots: Optional[Mapping[str, object]]) -> Mapping[str, tuple[str, ...]]:
|
||||
"""复制并冻结受保护根;任何非法根都在构造边界 fail-closed。"""
|
||||
if roots is None:
|
||||
return MappingProxyType({})
|
||||
if not isinstance(roots, Mapping):
|
||||
raise TypeError("reference_roots 必须是 recordId 到路径元组的映射")
|
||||
frozen = {}
|
||||
for record_id, raw_roots in roots.items():
|
||||
if isinstance(raw_roots, str):
|
||||
values = (raw_roots,)
|
||||
else:
|
||||
try:
|
||||
values = tuple(raw_roots)
|
||||
except TypeError as exc:
|
||||
raise TypeError("reference_roots 必须包含路径序列") from exc
|
||||
if not values:
|
||||
raise ValueError("reference_roots 不得包含空根")
|
||||
normalized = []
|
||||
for value in values:
|
||||
path = _normalize_repo_path(value)
|
||||
if path is None:
|
||||
raise ValueError("reference_roots 含非法路径")
|
||||
normalized.append(path)
|
||||
frozen[record_id] = tuple(normalized)
|
||||
return MappingProxyType(frozen)
|
||||
|
||||
|
||||
def _validate_reference_snapshot(
|
||||
files: Mapping[str, bytes], roots: Mapping[str, tuple[str, ...]]
|
||||
) -> None:
|
||||
"""校验快照与根索引的一致性,避免未覆盖文件回落到活目录。"""
|
||||
protected_roots = tuple(root for values in roots.values() for root in values)
|
||||
if files and not protected_roots:
|
||||
raise ValueError("reference_files 非空时 reference_roots 不能为空")
|
||||
for path in files:
|
||||
if not any(_path_is_within(path, root) for root in protected_roots):
|
||||
raise ValueError("reference_files 必须全部位于 reference_roots 内")
|
||||
|
||||
|
||||
def _protected_roots(session: "CheapSession") -> tuple[str, ...]:
|
||||
"""展平 session 的受保护根索引,供 read/list 共用同一边界判断。"""
|
||||
return tuple(root for roots in session.reference_roots.values() for root in roots)
|
||||
|
||||
|
||||
def _is_protected_path(session: "CheapSession", path: str) -> bool:
|
||||
"""判断规范化路径是否位于任一受保护根内。"""
|
||||
return any(_path_is_within(path, root) for root in _protected_roots(session))
|
||||
|
||||
|
||||
def _read_snapshot_bytes(content: bytes) -> str:
|
||||
"""按 cheap_run.read_file 的 200KB、UTF-8 忽略错误和 Toolkit 总上限返回文本。"""
|
||||
truncated = len(content) > _REFERENCE_READ_BYTES
|
||||
raw = content[:_REFERENCE_READ_BYTES] if truncated else content
|
||||
prefix = "[内容已截断]\n" if truncated else ""
|
||||
return (prefix + raw.decode("utf-8", "ignore"))[:_TOOL_RESULT_CAP]
|
||||
|
||||
|
||||
def _list_snapshot_directory(session: "CheapSession", path: str) -> str:
|
||||
"""从快照文件映射投影指定目录的直接子项,不观察活目录。"""
|
||||
entries = set()
|
||||
for file_path in session.reference_files:
|
||||
if not _is_protected_path(session, file_path) or not _path_is_within(file_path, path):
|
||||
continue
|
||||
suffix = file_path[len(path):].lstrip("/")
|
||||
if not suffix:
|
||||
continue
|
||||
first, separator, _ = suffix.partition("/")
|
||||
entries.add(first + "/" if separator else first)
|
||||
return "\n".join(sorted(entries))[:_TOOL_RESULT_CAP]
|
||||
|
||||
|
||||
@dataclass(init=False)
|
||||
class CheapSession:
|
||||
"""便宜档单 run 可变状态闭包(对 tier2 toolkit.py 的 Tier2Session,LittleJS 简化版)。
|
||||
|
||||
@ -27,6 +158,39 @@ class CheapSession:
|
||||
last_check: Optional[dict] = None
|
||||
last_build: Optional[dict] = None
|
||||
finished: Optional[dict] = None # finish 组装的产物摘要(None=未收敛,编排层据此判收敛)
|
||||
_reference_files: Mapping[str, bytes] = field(init=False, repr=False, compare=False)
|
||||
_reference_roots: Mapping[str, tuple[str, ...]] = field(init=False, repr=False, compare=False)
|
||||
|
||||
def __init__(
|
||||
self,
|
||||
game_id: str,
|
||||
last_check: Optional[dict] = None,
|
||||
last_build: Optional[dict] = None,
|
||||
finished: Optional[dict] = None,
|
||||
*,
|
||||
reference_files: Optional[Mapping[str, bytes]] = None,
|
||||
reference_roots: Optional[Mapping[str, object]] = None,
|
||||
) -> None:
|
||||
"""保留旧 session 参数顺序,并在边界复制冻结可信消费快照。"""
|
||||
self.game_id = game_id
|
||||
self.last_check = last_check
|
||||
self.last_build = last_build
|
||||
self.finished = finished
|
||||
frozen_files = _freeze_reference_files(reference_files)
|
||||
frozen_roots = _freeze_reference_roots(reference_roots)
|
||||
_validate_reference_snapshot(frozen_files, frozen_roots)
|
||||
object.__setattr__(self, "_reference_files", frozen_files)
|
||||
object.__setattr__(self, "_reference_roots", frozen_roots)
|
||||
|
||||
@property
|
||||
def reference_files(self) -> Mapping[str, bytes]:
|
||||
"""返回只读的仓内逻辑路径到 bytes 快照映射。"""
|
||||
return self._reference_files
|
||||
|
||||
@property
|
||||
def reference_roots(self) -> Mapping[str, tuple[str, ...]]:
|
||||
"""返回只读的 recordId 到受保护根路径元组映射。"""
|
||||
return self._reference_roots
|
||||
|
||||
|
||||
def build_toolkit(session: CheapSession, *, write_whitelist=None):
|
||||
@ -46,6 +210,14 @@ def build_toolkit(session: CheapSession, *, write_whitelist=None):
|
||||
Args:
|
||||
path: repo 相对路径,如 .agents/skills/littlejs-game-dev.md
|
||||
"""
|
||||
normalized = _normalize_repo_path(path)
|
||||
if normalized is not None and _is_protected_path(session, normalized):
|
||||
content = session.reference_files.get(normalized)
|
||||
if content is None:
|
||||
return _REFERENCE_DENY_MESSAGE
|
||||
return _read_snapshot_bytes(content)
|
||||
if _path_requires_rejection(path):
|
||||
return "ERROR: 路径必须是仓内相对路径"
|
||||
r = cheap_run.read_file(path)
|
||||
if not r["ok"]:
|
||||
return "ERROR: " + r["error"]
|
||||
@ -58,6 +230,11 @@ def build_toolkit(session: CheapSession, *, write_whitelist=None):
|
||||
Args:
|
||||
path: repo 相对路径。
|
||||
"""
|
||||
normalized = _normalize_repo_path(path)
|
||||
if normalized is not None and _is_protected_path(session, normalized):
|
||||
return _list_snapshot_directory(session, normalized)
|
||||
if _path_requires_rejection(path):
|
||||
return "ERROR: 路径必须是仓内相对路径"
|
||||
r = cheap_run.list_dir(path)
|
||||
if not r["ok"]:
|
||||
return "ERROR: " + r["error"]
|
||||
|
||||
File diff suppressed because it is too large
Load Diff
284
cheap-worker/full_gate.py
Normal file
284
cheap-worker/full_gate.py
Normal file
@ -0,0 +1,284 @@
|
||||
"""full_gate.py — 生成线验收 full_gate 集成 runner(W-AXIS 收口 R1)。
|
||||
|
||||
把六个已成熟的子门部件串成单一 full_gate decision,只消费子门既有产物/函数,不重写任何子门逻辑:
|
||||
|
||||
① 九门机械预筛 ← Node tools.mjs check + play.cdp.cjs 驱动器产出的 verdict(pass/guards)
|
||||
② 视觉地板 ← cheap_verify.judge_gameplay_floor 的裁决(parse_floor_judgment 结构:accepted/degraded)
|
||||
③ playtest/3 真玩 ← cheap_verify.run_acceptance_v3 封存 payload(双 Judge + rollGuard + merge + finalPostguard)
|
||||
④ prompt 四门 ← eval_gate.py 真模型闸台账记录(gate1 schema / gate2 成功率 / gate3 回归 / gate4 成本延迟 + gate5 稳定)
|
||||
⑤ schema 语义 ← contracts/play-loop/validate.py(经 cheap_verify.validate_acceptance_v3_payload 调 canonical 校验)
|
||||
⑥ W-GOLD-LIVE 对账 ← cheap_verify.reconcile_v3_reference_asset_consumption(参照资产消费六项闸)
|
||||
|
||||
降级口径(设计档 §3.2 结果四态 + §5.2 失败归因,焊死):
|
||||
· 全过 → outcome=accept、pass=True(唯一可接受路径);
|
||||
· 任一子门 tester_error / 缺产物 / degraded / 自相矛盾 → 整体 tester_error(仪器异常不得伪装成 accept,
|
||||
也不得伪装成 gameplay reject——§5.2「Actor/runner/Judge/schema/图像/环境异常 100% 归 tester_error」);
|
||||
· 任一子门 reject → 整体 reject(硬证已证的产物缺陷);
|
||||
· 任一子门 inconclusive → 整体 inconclusive(合法证据不足或矛盾);
|
||||
· 优先级 tester_error > reject > inconclusive > accept:仪器异常时证据不可信,不能 claiming 已证缺陷。
|
||||
|
||||
本步为纯本地代码 + 罐头单测;真跑(真模型真浏览器)留 mini-desktop,经 run_full_gate_live 编排。
|
||||
"""
|
||||
|
||||
import sys
|
||||
from pathlib import Path
|
||||
|
||||
# 与 cheap-worker 其余模块同惯例:本目录直挂 sys.path,import cheap_verify 复用子门函数。
|
||||
sys.path.insert(0, str(Path(__file__).resolve().parent))
|
||||
|
||||
import cheap_verify as V # noqa: E402 子门函数单一来源:地板判定/v3 编排/canonical 校验/金标对账
|
||||
|
||||
# 设计档 §3.2 结果四态;full_gate 与子门共用同一取值集,未知取值一律 fail-closed 归 tester_error。
|
||||
_OUTCOMES = ("accept", "reject", "inconclusive", "tester_error")
|
||||
|
||||
# 子门状态合并优先级:数值越小越严重。仪器异常(tester_error)压过已证缺陷(reject)压过证据不足(inconclusive)。
|
||||
_SEVERITY = {"tester_error": 0, "reject": 1, "inconclusive": 2, "pass": 3}
|
||||
|
||||
|
||||
# ────────────────────────── 六个子门消费函数(只消费产物,不重写逻辑)──────────────────────────
|
||||
|
||||
def _eval_mechanical_nine(verdict) -> dict:
|
||||
"""① 九门机械预筛:消费 Node harness verdict(A..I guards + pass AND)。
|
||||
|
||||
缺产物 → tester_error(无机械证据,fail-closed);guards 与 pass 自相矛盾 → tester_error(harness 异常);
|
||||
明确未过 → reject(结构性/装载/渲染/接线机械缺陷)。
|
||||
"""
|
||||
if not isinstance(verdict, dict):
|
||||
return {"state": "tester_error", "reason": "九门 verdict 产物缺失,无机械预筛证据", "failedGates": []}
|
||||
# check 级结构失败(如 src/ 不存在):verdict.ok=False 携带 errors,属机械缺陷。
|
||||
if verdict.get("ok") is False:
|
||||
errors = [str(e) for e in (verdict.get("errors") or [])][:6]
|
||||
return {"state": "reject", "reason": "机械结构检查未过:" + ";".join(errors), "failedGates": []}
|
||||
guards = verdict.get("guards") if isinstance(verdict.get("guards"), dict) else {}
|
||||
failed = sorted(name for name, g in guards.items() if isinstance(g, dict) and g.get("pass") is False)
|
||||
passed = verdict.get("pass")
|
||||
if passed is True and failed:
|
||||
# pass=True 却存在 pass=False 的门:harness 自报与逐门证据矛盾 → 仪器异常,不信任一自报。
|
||||
return {"state": "tester_error",
|
||||
"reason": f"verdict.pass=True 与 guards 矛盾(未过门:{failed})", "failedGates": failed}
|
||||
if passed is True:
|
||||
return {"state": "pass", "reason": "九门机械预筛全绿", "failedGates": []}
|
||||
if passed is False:
|
||||
return {"state": "reject", "reason": f"九门机械预筛未过({','.join(failed) or 'pass=False'})",
|
||||
"failedGates": failed}
|
||||
return {"state": "tester_error", "reason": "verdict.pass 缺失或非布尔,机械预筛无结论", "failedGates": failed}
|
||||
|
||||
|
||||
def _eval_visual_floor(floor_judgment) -> dict:
|
||||
"""② 视觉地板:消费独立模型玩法地板判定(judge_gameplay_floor / parse_floor_judgment 结构)。
|
||||
|
||||
degraded(评不出/无证据/调用失败)按 §5.2 归仪器异常 tester_error——不伪装 accept,也不伪装 gameplay reject;
|
||||
accepted=False 且非 degraded → reject(独立模型判定的 broken/hollow/off_brief)。
|
||||
"""
|
||||
if not isinstance(floor_judgment, dict):
|
||||
return {"state": "tester_error", "reason": "视觉地板判定产物缺失", "rejectClasses": []}
|
||||
if floor_judgment.get("degraded") is True:
|
||||
return {"state": "tester_error",
|
||||
"reason": f"视觉地板降级(fail-closed):{floor_judgment.get('reason') or '未知'}", "rejectClasses": []}
|
||||
accepted = floor_judgment.get("accepted")
|
||||
if accepted is True:
|
||||
return {"state": "pass", "reason": "视觉地板接受", "rejectClasses": []}
|
||||
if accepted is False:
|
||||
classes = [str(c) for c in (floor_judgment.get("rejectClasses") or [])]
|
||||
return {"state": "reject", "reason": f"视觉地板拒绝({','.join(classes) or '地板未过'})",
|
||||
"rejectClasses": classes}
|
||||
return {"state": "tester_error", "reason": "视觉地板判定 accepted 字段缺失或非布尔", "rejectClasses": []}
|
||||
|
||||
|
||||
def _eval_playtest_v3(v3_payload) -> dict:
|
||||
"""③ playtest/3 真玩:消费 run_acceptance_v3 封存 payload 的 outcome/decision/finalPostguard。
|
||||
|
||||
不复算子门内部逻辑,但做零信任自洽复核:accepted=True 必须与 outcome=accept 且 finalPostguard.pass=True
|
||||
同真,任一自相矛盾 → tester_error(封存产物不一致 = 仪器异常)。shadow 模式 accepted 仅供校准,
|
||||
publishFrozen 原样带出,发布判定留给 is_v3_publishable。
|
||||
"""
|
||||
if not isinstance(v3_payload, dict):
|
||||
return {"state": "tester_error", "reason": "playtest/3 封存产物缺失", "outcome": None}
|
||||
outcome = v3_payload.get("outcome")
|
||||
decision = v3_payload.get("decision") if isinstance(v3_payload.get("decision"), dict) else {}
|
||||
final_guard = v3_payload.get("finalPostguard") if isinstance(v3_payload.get("finalPostguard"), dict) else {}
|
||||
compatibility = v3_payload.get("compatibility") if isinstance(v3_payload.get("compatibility"), dict) else {}
|
||||
detail = {
|
||||
"outcome": outcome, "accepted": decision.get("accepted"),
|
||||
"acceptanceMode": v3_payload.get("acceptanceMode"),
|
||||
"publishFrozen": decision.get("publishFrozen"),
|
||||
"rescuedByRoll": decision.get("rescuedByRoll"),
|
||||
"failure": decision.get("failure"),
|
||||
}
|
||||
if outcome not in _OUTCOMES:
|
||||
return {"state": "tester_error", "reason": f"playtest/3 outcome 未知:{outcome!r}", **detail}
|
||||
# 零信任自洽复核:三个写权字段必须同真,矛盾即封存产物异常。
|
||||
if decision.get("accepted") is True and (outcome != "accept" or final_guard.get("pass") is not True):
|
||||
return {"state": "tester_error",
|
||||
"reason": f"decision.accepted=True 与 outcome={outcome}/finalPostguard.pass="
|
||||
f"{final_guard.get('pass')} 矛盾", **detail}
|
||||
if outcome == "accept":
|
||||
if decision.get("accepted") is not True or final_guard.get("pass") is not True:
|
||||
return {"state": "tester_error",
|
||||
"reason": "outcome=accept 但 decision/finalPostguard 未同时确认", **detail}
|
||||
return {"state": "pass", "reason": "playtest/3 真玩接受(硬证完整 + 双 Judge 共识 + finalPostguard 全绿)",
|
||||
"compatibilityAccepted": compatibility.get("accepted"), **detail}
|
||||
if outcome == "reject":
|
||||
return {"state": "reject", "reason": "playtest/3 真玩拒绝(硬证已证可复现产物缺陷)", **detail}
|
||||
if outcome == "inconclusive":
|
||||
return {"state": "inconclusive", "reason": "playtest/3 证据不足或矛盾", **detail}
|
||||
return {"state": "tester_error", "reason": "playtest/3 仪器异常(Actor/runner/Judge/环境)", **detail}
|
||||
|
||||
|
||||
def _eval_prompt_eval(prompt_eval_record) -> dict:
|
||||
"""④ prompt 四门:消费 eval_gate.py 真模型闸台账记录(单条 dict 或 Actor/Judge 多条 list)。
|
||||
|
||||
每条记录必须 infrastructureComplete 且 allGreen(= gate1 schema ∧ gate2 成功率 ∧ gate3 回归 ∧ gate4 成本延迟
|
||||
∧ gate5 稳定 全绿)。校准未过 = Actor/Judge 判读不可信,按 §5.2「Judge 异常 100% tester_error」归仪器异常,
|
||||
绝不放行也绝不归 gameplay。
|
||||
"""
|
||||
if prompt_eval_record is None:
|
||||
return {"state": "tester_error", "reason": "prompt 四门台账记录缺失(校准闸未跑)", "failedRecords": []}
|
||||
records = prompt_eval_record if isinstance(prompt_eval_record, list) else [prompt_eval_record]
|
||||
if not records:
|
||||
return {"state": "tester_error", "reason": "prompt 四门台账记录为空", "failedRecords": []}
|
||||
failed = []
|
||||
for i, record in enumerate(records):
|
||||
if not isinstance(record, dict):
|
||||
failed.append({"index": i, "promptId": None, "reason": "记录不是对象"})
|
||||
continue
|
||||
prompt_id = record.get("promptId") or f"#{i}"
|
||||
if record.get("infrastructureComplete") is not True:
|
||||
failed.append({"index": i, "promptId": prompt_id, "reason": "基础设施不完整(图像/审计/请求哈希不可信)"})
|
||||
continue
|
||||
# allGreen 缺失时退回逐门 AND(gate1..gate5),两者皆无 → fail-closed。
|
||||
all_green = record.get("allGreen")
|
||||
if all_green is None:
|
||||
gate_keys = ("gate1_schema", "gate2_success", "gate3_regression", "gate4_cost_latency", "gate5_stability")
|
||||
gates = {key: record.get(key) for key in gate_keys}
|
||||
if any(value is None for value in gates.values()):
|
||||
failed.append({"index": i, "promptId": prompt_id, "reason": "缺 allGreen 且逐门结果不全"})
|
||||
continue
|
||||
all_green = all(value is True for value in gates.values())
|
||||
if all_green is not True:
|
||||
red = [key for key in ("gate1_schema", "gate2_success", "gate3_regression",
|
||||
"gate4_cost_latency", "gate5_stability") if record.get(key) is False]
|
||||
failed.append({"index": i, "promptId": prompt_id, "reason": f"闸门红:{','.join(red) or 'allGreen=False'}"})
|
||||
if failed:
|
||||
return {"state": "tester_error", "reason": f"prompt 四门未全绿({len(failed)}/{len(records)} 条)",
|
||||
"failedRecords": failed}
|
||||
return {"state": "pass", "reason": f"prompt 四门全绿({len(records)} 条记录)", "failedRecords": []}
|
||||
|
||||
|
||||
def _eval_schema_semantics(v3_payload, schema_validator) -> dict:
|
||||
"""⑤ schema 语义:对 playtest/3 封存产物跑 canonical validate.py(schema + 语义)。
|
||||
|
||||
经 cheap_verify.validate_acceptance_v3_payload 调权威校验器(运行时与测试共用同一规则);
|
||||
任何错误按 §5.2「schema 异常 100% tester_error」归仪器异常。
|
||||
"""
|
||||
if not isinstance(v3_payload, dict):
|
||||
return {"state": "tester_error", "reason": "无 playtest/3 产物可校验", "errors": []}
|
||||
try:
|
||||
errors = [str(e) for e in (schema_validator(v3_payload) or [])]
|
||||
except Exception as exc: # noqa: BLE001 —— 校验器自身异常必须 fail-closed 归仪器异常
|
||||
return {"state": "tester_error", "reason": f"canonical 校验器异常:{type(exc).__name__}: {exc}", "errors": []}
|
||||
if errors:
|
||||
return {"state": "tester_error", "reason": f"playtest/3 schema/语义校验失败({len(errors)} 项)",
|
||||
"errors": errors[:8]}
|
||||
return {"state": "pass", "reason": "canonical schema + 语义校验通过", "errors": []}
|
||||
|
||||
|
||||
def _eval_gold_reconcile(consumptions, consumer_ref, gold_registry, gold_reconciler) -> dict:
|
||||
"""⑥ W-GOLD-LIVE 消费对账:对 identity 声明的参照资产消费跑六项闸(存在/激活/role/consumerRef/版本/缺维度)。
|
||||
|
||||
未声明消费 ≡ 旧路径,vacuous 放行(与 _v3_check_declared_reference_assets 语义一致);
|
||||
声明消费而对账失败 → reject(设计语义:声明消费无 active 匹配 = verified reject)。
|
||||
"""
|
||||
items = list(consumptions or [])
|
||||
if not items:
|
||||
return {"state": "pass", "reason": "未声明参照资产消费(旧路径,对账 vacuous 通过)",
|
||||
"vacuous": True, "consumed": [], "errors": []}
|
||||
try:
|
||||
reconciled = gold_reconciler(items, consumer_ref=consumer_ref, registry=gold_registry)
|
||||
except Exception as exc: # noqa: BLE001 —— 注册表不可读按 fail-closed 拒绝消费
|
||||
return {"state": "reject", "reason": f"参照资产消费对账失败:{type(exc).__name__}: {exc}",
|
||||
"vacuous": False, "consumed": [], "errors": [str(exc)]}
|
||||
if reconciled.get("ok"):
|
||||
return {"state": "pass", "reason": f"参照资产消费对账通过({len(reconciled.get('consumed') or [])} 条)",
|
||||
"vacuous": False, "consumed": reconciled.get("consumed") or [], "errors": []}
|
||||
return {"state": "reject", "reason": "参照资产消费对账拒绝:" + ";".join((reconciled.get("errors") or [])[:4]),
|
||||
"vacuous": False, "consumed": [], "errors": reconciled.get("errors") or []}
|
||||
|
||||
|
||||
# ────────────────────────── 单一 decision 组合(确定性,零模型)──────────────────────────
|
||||
|
||||
def _combine(gates: dict) -> dict:
|
||||
"""按严重度优先级合并六个子门状态为单一 full_gate decision(tester_error > reject > inconclusive > accept)。"""
|
||||
states = [g["state"] for g in gates.values()]
|
||||
worst = min(states, key=lambda s: _SEVERITY.get(s, -1))
|
||||
if worst not in _SEVERITY:
|
||||
# 出现未知状态本身即仪器异常,绝不放行。
|
||||
worst = "tester_error"
|
||||
outcome = "accept" if worst == "pass" else worst
|
||||
reasons = [f"[{name}] {g['reason']}" for name, g in gates.items() if g["state"] != "pass"]
|
||||
return {"pass": worst == "pass", "outcome": outcome, "reasons": reasons}
|
||||
|
||||
|
||||
def run_full_gate(*, verdict=None, floor_judgment=None, v3_payload=None,
|
||||
prompt_eval_record=None, reference_consumptions=(),
|
||||
consumer_ref=None, gold_registry=None,
|
||||
schema_validator=None, gold_reconciler=None) -> dict:
|
||||
"""full_gate 纯函数核心:消费六个子门既有产物,产出单一 decision(确定性、零模型、零 I/O)。
|
||||
|
||||
Args:
|
||||
verdict: Node 九门产物 {ok?, pass, guards, errors?};None=未跑 → fail-closed tester_error。
|
||||
floor_judgment: judge_gameplay_floor 裁决 {accepted, degraded, rejectClasses, reason?}。
|
||||
v3_payload: run_acceptance_v3 封存 payload(outcome/decision/finalPostguard/compatibility)。
|
||||
prompt_eval_record: eval_gate 台账记录 dict 或 list(Actor + Judge A/B 多条)。
|
||||
reference_consumptions: identity 声明的参照资产消费(recordId 字符串或对象列表);空=vacuous 放行。
|
||||
consumer_ref / gold_registry: 金标对账的消费方身份与注册表(None → 默认注册表)。
|
||||
schema_validator / gold_reconciler: 可注入的子门函数(默认绑 cheap_verify 真函数;单测注入罐头)。
|
||||
|
||||
Returns:
|
||||
{pass, outcome, publishable, gates: {六门各自 state/reason/细节}, reasons: [未过原因]}
|
||||
"""
|
||||
validator = schema_validator if schema_validator is not None else V.validate_acceptance_v3_payload
|
||||
reconciler = gold_reconciler if gold_reconciler is not None else V.reconcile_v3_reference_asset_consumption
|
||||
|
||||
gates = {
|
||||
"mechanicalNine": _eval_mechanical_nine(verdict),
|
||||
"visualFloor": _eval_visual_floor(floor_judgment),
|
||||
"playtestV3": _eval_playtest_v3(v3_payload),
|
||||
"promptEval": _eval_prompt_eval(prompt_eval_record),
|
||||
"schemaSemantics": _eval_schema_semantics(v3_payload, validator),
|
||||
"goldReconcile": _eval_gold_reconcile(reference_consumptions, consumer_ref, gold_registry, reconciler),
|
||||
}
|
||||
decision = _combine(gates)
|
||||
|
||||
# 发布谓词只在全过时才有意义:active v3 + 未冻结 + 兼容层同真(is_v3_publishable 的字段口径,消费不重写)。
|
||||
publishable = False
|
||||
if decision["pass"] and isinstance(v3_payload, dict):
|
||||
decision_node = v3_payload.get("decision") if isinstance(v3_payload.get("decision"), dict) else {}
|
||||
compatibility = v3_payload.get("compatibility") if isinstance(v3_payload.get("compatibility"), dict) else {}
|
||||
publishable = bool(
|
||||
v3_payload.get("acceptanceMode") == "v3" and decision_node.get("publishFrozen") is False
|
||||
and compatibility.get("accepted") is True and compatibility.get("ok") is True
|
||||
and compatibility.get("publishFrozen") is False)
|
||||
decision["publishable"] = publishable
|
||||
decision["gates"] = gates
|
||||
return decision
|
||||
|
||||
|
||||
async def run_full_gate_live(*, game_id: str, brief: str, verdict: dict, v3_request: dict,
|
||||
floor_model=None, floor_model_name: str = None,
|
||||
prompt_eval_record=None, reference_consumptions=(),
|
||||
consumer_ref=None, gold_registry=None) -> dict:
|
||||
"""真编排入口(mini-desktop 真跑用):调视觉地板判定 + run_acceptance_v3 真玩,再交纯函数核心出单一 decision。
|
||||
|
||||
本函数是唯一会触达真模型的编排层;九门 verdict 与 prompt 四门台账由上游生成/CI 产物传入(消费既有产物)。
|
||||
本地单测只覆盖 run_full_gate 纯核心,不跑本函数。
|
||||
"""
|
||||
# ② 视觉地板:独立模型看真玩截图/日志(fail-closed 语义由 judge_gameplay_floor 内部保证)。
|
||||
floor_judgment = await V.judge_gameplay_floor(
|
||||
game_id, brief=brief, model=floor_model, model_name=floor_model_name)
|
||||
# ③ playtest/3 真玩:唯一 v3 编排入口(含四门投影/双 Judge/二掷/merge/finalPostguard/幂等封存)。
|
||||
v3_payload = await V.run_acceptance_v3(v3_request)
|
||||
return run_full_gate(verdict=verdict, floor_judgment=floor_judgment, v3_payload=v3_payload,
|
||||
prompt_eval_record=prompt_eval_record,
|
||||
reference_consumptions=reference_consumptions,
|
||||
consumer_ref=consumer_ref, gold_registry=gold_registry)
|
||||
812
cheap-worker/reference_asset_gate.py
Normal file
812
cheap-worker/reference_asset_gate.py
Normal file
@ -0,0 +1,812 @@
|
||||
"""参照资产 v2 的受信 fd 消费门与双身份验证器。"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import hashlib
|
||||
import json
|
||||
import os
|
||||
import re
|
||||
import unicodedata
|
||||
from collections.abc import Mapping, Sequence
|
||||
from dataclasses import dataclass
|
||||
from pathlib import Path
|
||||
from types import MappingProxyType
|
||||
|
||||
try:
|
||||
import artifact_snapshot
|
||||
except ImportError: # pragma: no cover - 直接以 cheap-worker 目录为 cwd 时使用绝对导入
|
||||
from . import artifact_snapshot
|
||||
|
||||
|
||||
POLICY_ID = "survivor-gold-v1"
|
||||
POLICY_RECORD_ID = "gac-shanhai-xingji"
|
||||
POLICY_ROLE = "game_content_gold"
|
||||
POLICY_CONSUMER_REF = "generation-runtime@reference-assets/2"
|
||||
POLICY_MODE = "frozen_preflight"
|
||||
VERIFIER_VERSION = "reference-asset-verifier/1.0.0"
|
||||
TRUSTED_ROOT_ID = "wanxiang-reference-assets-root-v1"
|
||||
SNAPSHOT_HASH_DOMAIN = b"reference-asset-consumption-snapshot/1\n"
|
||||
|
||||
MAX_MANIFEST_BYTES = 1 * 1024 * 1024
|
||||
MAX_MANIFEST_ENTRIES = 512
|
||||
MAX_FILE_BYTES = 16 * 1024 * 1024
|
||||
MAX_RECORD_BYTES = 64 * 1024 * 1024
|
||||
MAX_TOTAL_BYTES = 128 * 1024 * 1024
|
||||
|
||||
_SHA256_RE = re.compile(r"^[0-9a-f]{64}$")
|
||||
_IDENTIFIER_RE = re.compile(r"^[A-Za-z0-9_][A-Za-z0-9._-]*$")
|
||||
_VERIFIER_RE = re.compile(r"^[A-Za-z][A-Za-z0-9._-]*/[0-9]+\.[0-9]+\.[0-9]+$")
|
||||
_TRUSTED_ROOT_ID_RE = re.compile(r"^[A-Za-z0-9_][A-Za-z0-9._:-]*$")
|
||||
_SIGNED_AT_RE = re.compile(
|
||||
r"^\d{4}-\d{2}-\d{2}(T\d{2}:\d{2}(:\d{2})?(Z|[+-]\d{2}:?\d{2})?)?$"
|
||||
)
|
||||
|
||||
_ALLOWED_ROLES = {
|
||||
"harness_fixture",
|
||||
"prompt_eval_gold",
|
||||
"generation_exemplar",
|
||||
"game_content_gold",
|
||||
}
|
||||
_ALLOWED_LIFECYCLE = {"candidate", "migration_pending", "active", "retired"}
|
||||
_RECORD_KEYS = {
|
||||
"schemaVersion",
|
||||
"recordId",
|
||||
"role",
|
||||
"lifecycleStatus",
|
||||
"assetRef",
|
||||
"assetVersion",
|
||||
"artifactHash",
|
||||
"consumerRef",
|
||||
"designRef",
|
||||
"evidenceRefs",
|
||||
"signedBy",
|
||||
"signedAt",
|
||||
"artifactRef",
|
||||
"consumptionManifestRef",
|
||||
"consumptionManifestHash",
|
||||
}
|
||||
_RELEASE_KEYS = {
|
||||
"schemaVersion",
|
||||
"releaseId",
|
||||
"registryRef",
|
||||
"registryHash",
|
||||
"policyRef",
|
||||
"policyHash",
|
||||
"verifierVersion",
|
||||
"trustedRootId",
|
||||
}
|
||||
_POLICY_KEYS = {
|
||||
"schemaVersion",
|
||||
"policyId",
|
||||
"recordId",
|
||||
"role",
|
||||
"consumerRef",
|
||||
"route",
|
||||
"autoSelect",
|
||||
"mode",
|
||||
}
|
||||
_MANIFEST_KEYS = {"schemaVersion", "manifestId", "canonicalization", "entries"}
|
||||
_MANIFEST_ENTRY_KEYS = {"path", "size", "sha256"}
|
||||
|
||||
|
||||
class ReferenceAssetGateError(ValueError):
|
||||
"""稳定的可信消费错误;异常正文只含 code、recordId 和逻辑路径。"""
|
||||
|
||||
def __init__(self, code: str, record_id: str | None = None, logical_path: str | None = None) -> None:
|
||||
self.code = code
|
||||
self.record_id = record_id
|
||||
self.logical_path = logical_path
|
||||
fields = [f"code={code}"]
|
||||
if record_id is not None:
|
||||
fields.append(f"recordId={record_id}")
|
||||
if logical_path is not None:
|
||||
fields.append(f"path={logical_path}")
|
||||
super().__init__(" ".join(fields))
|
||||
|
||||
|
||||
# 为调用方保留两个直观别名,实际异常类型只有一套稳定字段。
|
||||
ReferenceAssetError = ReferenceAssetGateError
|
||||
ReferenceAssetVerificationError = ReferenceAssetGateError
|
||||
|
||||
|
||||
@dataclass(frozen=True)
|
||||
class VerifiedReferenceAssets:
|
||||
"""验证成功后的只读消费闭包;失败时不会构造或返回该对象。"""
|
||||
|
||||
constraint_records: tuple[Mapping[str, object], ...]
|
||||
files: Mapping[str, bytes]
|
||||
receipts: tuple[Mapping[str, object], ...]
|
||||
snapshot_hash: str
|
||||
reference_roots: Mapping[str, tuple[str, ...]]
|
||||
|
||||
@property
|
||||
def records(self) -> tuple[Mapping[str, object], ...]:
|
||||
"""兼容调用方使用 records 读取已冻结的约束记录。"""
|
||||
return self.constraint_records
|
||||
|
||||
@property
|
||||
def canonical_snapshot_hash(self) -> str:
|
||||
"""返回消费快照 canonical hash 的显式别名。"""
|
||||
return self.snapshot_hash
|
||||
|
||||
@property
|
||||
def reference_files(self) -> Mapping[str, bytes]:
|
||||
"""返回 Task 4 Toolkit 使用的只读文件映射。"""
|
||||
return self.files
|
||||
|
||||
@property
|
||||
def protected_roots(self) -> Mapping[str, tuple[str, ...]]:
|
||||
"""返回按 recordId 索引的受保护根集合。"""
|
||||
return self.reference_roots
|
||||
|
||||
|
||||
@dataclass(frozen=True)
|
||||
class _StagedReferenceAssets:
|
||||
"""单条记录的内部暂存结果;batch 全部成功前不暴露公开结果对象。"""
|
||||
|
||||
record: Mapping[str, object]
|
||||
files: Mapping[str, bytes]
|
||||
receipt: Mapping[str, object]
|
||||
roots: tuple[str, ...]
|
||||
|
||||
|
||||
def _freeze(value):
|
||||
"""递归冻结记录、回执和根索引,防止调用方改写验证结果。"""
|
||||
if isinstance(value, Mapping):
|
||||
return MappingProxyType({key: _freeze(item) for key, item in value.items()})
|
||||
if isinstance(value, (list, tuple)):
|
||||
return tuple(_freeze(item) for item in value)
|
||||
return value
|
||||
|
||||
|
||||
def to_json_value(value):
|
||||
"""把只读验证结果显式复制成普通 JSON 容器,供落盘和 schema 校验边界使用。"""
|
||||
if isinstance(value, Mapping):
|
||||
return {key: to_json_value(item) for key, item in value.items()}
|
||||
if isinstance(value, tuple):
|
||||
return [to_json_value(item) for item in value]
|
||||
return value
|
||||
|
||||
|
||||
def canonical_json_bytes(value: object) -> bytes:
|
||||
"""生成 manifest 使用的 UTF-8 canonical JSON 原始字节。"""
|
||||
return json.dumps(
|
||||
value,
|
||||
ensure_ascii=False,
|
||||
sort_keys=True,
|
||||
separators=(",", ":"),
|
||||
allow_nan=False,
|
||||
).encode("utf-8")
|
||||
|
||||
|
||||
def snapshot_hash(files: Mapping[str, bytes]) -> str:
|
||||
"""按路径 UTF-8 字节序计算消费快照 hash 向量。"""
|
||||
return artifact_snapshot.consumption_snapshot_hash(files)
|
||||
|
||||
|
||||
def canonical_snapshot_hash(files: Mapping[str, bytes]) -> str:
|
||||
"""snapshot_hash 的公开语义别名,供跨语言向量测试调用。"""
|
||||
return snapshot_hash(files)
|
||||
|
||||
|
||||
def _fail(code: str, record_id: str | None = None, logical_path: str | None = None):
|
||||
"""统一抛出稳定异常,禁止把底层系统错误和输入内容向外传播。"""
|
||||
raise ReferenceAssetGateError(code, record_id, logical_path)
|
||||
|
||||
|
||||
def _path(value, *, record_id: str | None = None) -> str:
|
||||
"""校验仓根相对 NFC POSIX 路径;越界与其它路径语法错误分码。"""
|
||||
if isinstance(value, Path):
|
||||
value = value.as_posix()
|
||||
if not isinstance(value, str) or not value or len(value) > 1024 or "\x00" in value:
|
||||
_fail("reference_path_invalid", record_id, "<path>")
|
||||
if value.startswith("/"):
|
||||
_fail("reference_path_escape", record_id, "<absolute>")
|
||||
if "\\" in value or unicodedata.normalize("NFC", value) != value:
|
||||
_fail("reference_path_invalid", record_id, "<path>")
|
||||
parts = value.split("/")
|
||||
if any(part == ".." for part in parts):
|
||||
_fail("reference_path_escape", record_id, "<path>")
|
||||
if any(part in ("", ".") for part in parts):
|
||||
_fail("reference_path_invalid", record_id, "<path>")
|
||||
if any(ord(char) < 0x20 or ord(char) == 0x7F for char in value):
|
||||
_fail("reference_path_invalid", record_id, "<path>")
|
||||
return value
|
||||
|
||||
|
||||
def _keys(value, allowed: set[str], *, record_id: str | None = None, code: str = "reference_registry_invalid") -> None:
|
||||
"""执行 strict object key 检查,拒绝未知字段和非对象。"""
|
||||
if not isinstance(value, dict) or set(value) != allowed:
|
||||
_fail(code, record_id)
|
||||
|
||||
|
||||
def _keys_optional(value, required: set[str], optional: set[str], *, record_id: str | None = None) -> None:
|
||||
"""执行带 optional 字段的 strict object key 检查。"""
|
||||
if not isinstance(value, dict) or not required.issubset(value) or set(value) - required - optional:
|
||||
_fail("reference_registry_invalid", record_id)
|
||||
|
||||
|
||||
def _is_sha256(value) -> bool:
|
||||
"""判断是否为小写 64 位 SHA-256 字符串。"""
|
||||
return isinstance(value, str) and _SHA256_RE.fullmatch(value) is not None
|
||||
|
||||
|
||||
def _is_identifier(value) -> bool:
|
||||
"""判断 schema 中的稳定标识符。"""
|
||||
return isinstance(value, str) and _IDENTIFIER_RE.fullmatch(value) is not None
|
||||
|
||||
|
||||
def _load_json(raw: bytes, *, code: str, record_id: str | None, logical_path: str):
|
||||
"""只解析可信 fd 已读取的 UTF-8 JSON,并拒绝重复 object key。"""
|
||||
def pairs(pairs_list):
|
||||
result = {}
|
||||
for key, value in pairs_list:
|
||||
if key in result:
|
||||
raise ValueError("duplicate-key")
|
||||
result[key] = value
|
||||
return result
|
||||
|
||||
try:
|
||||
text = raw.decode("utf-8")
|
||||
value = json.loads(text, object_pairs_hook=pairs, parse_constant=lambda _: (_ for _ in ()).throw(ValueError()))
|
||||
except (UnicodeDecodeError, json.JSONDecodeError, ValueError):
|
||||
_fail(code, record_id, logical_path)
|
||||
if not isinstance(value, dict):
|
||||
_fail(code, record_id, logical_path)
|
||||
return value
|
||||
|
||||
|
||||
def _read_files(root_fd: int, paths: Sequence[str], *, limits: Mapping[str, int], record_id: str | None,
|
||||
missing_code: str | None = None) -> dict[str, bytes]:
|
||||
"""通过 artifact_snapshot 的受信选择性读取器取回文件,统一异常类型。"""
|
||||
try:
|
||||
snapshot = artifact_snapshot.capture_selected_files(root_fd, paths, limits)
|
||||
except artifact_snapshot.ArtifactSnapshotError as exc:
|
||||
code = missing_code if missing_code is not None and exc.code == "reference_missing" else exc.code
|
||||
_fail(code, record_id, exc.logical_path)
|
||||
except (OSError, ValueError, TypeError):
|
||||
_fail("reference_unreadable", record_id)
|
||||
return dict(snapshot.files)
|
||||
|
||||
|
||||
def _open_trusted_root(trusted_root) -> int:
|
||||
"""只打开可信根一次,之后所有引用均通过该 fd 解析。"""
|
||||
try:
|
||||
root_fd, _ = artifact_snapshot._selected_root_fd(trusted_root)
|
||||
return root_fd
|
||||
except artifact_snapshot.ArtifactSnapshotError as exc:
|
||||
_fail(exc.code, logical_path=exc.logical_path)
|
||||
except (OSError, TypeError, ValueError):
|
||||
_fail("reference_unreadable", logical_path="<root>")
|
||||
|
||||
|
||||
def _read_json_file(root_fd: int, logical_path: str, *, record_id: str | None, limits: Mapping[str, int],
|
||||
missing_code: str | None = None) -> tuple[dict, bytes, str]:
|
||||
"""从同一可信根读取 JSON,并返回对象、原始字节和原始 SHA-256。"""
|
||||
logical_path = _path(logical_path, record_id=record_id)
|
||||
raw = _read_files(
|
||||
root_fd,
|
||||
[logical_path],
|
||||
limits=limits,
|
||||
record_id=record_id,
|
||||
missing_code=missing_code,
|
||||
)[logical_path]
|
||||
value = _load_json(raw, code="reference_registry_invalid", record_id=record_id, logical_path=logical_path)
|
||||
return value, raw, hashlib.sha256(raw).hexdigest()
|
||||
|
||||
|
||||
def _validate_release(release: dict, *, record_id: str | None = None) -> None:
|
||||
"""校验 release/1 的 strict 结构和引用字段。"""
|
||||
_keys(release, _RELEASE_KEYS, record_id=record_id)
|
||||
if release.get("schemaVersion") != "ReferenceAssetRelease/1":
|
||||
_fail("reference_registry_invalid", record_id)
|
||||
if not _is_identifier(release.get("releaseId")):
|
||||
_fail("reference_registry_invalid", record_id)
|
||||
for field in ("registryRef", "policyRef"):
|
||||
_path(release.get(field), record_id=record_id)
|
||||
for field in ("registryHash", "policyHash"):
|
||||
if not _is_sha256(release.get(field)):
|
||||
_fail("reference_registry_invalid", record_id)
|
||||
if release.get("verifierVersion") != VERIFIER_VERSION:
|
||||
_fail("reference_registry_invalid", record_id)
|
||||
if release.get("trustedRootId") != TRUSTED_ROOT_ID:
|
||||
_fail("reference_registry_invalid", record_id)
|
||||
|
||||
|
||||
def _validate_nullable_string(value) -> bool:
|
||||
"""检查允许 null 的非空字符串字段。"""
|
||||
return value is None or (isinstance(value, str) and bool(value))
|
||||
|
||||
|
||||
def _validate_record(record: dict, *, record_id: str | None = None) -> None:
|
||||
"""校验 Record/2 字段、生命周期条件和所有仓内路径。"""
|
||||
_keys_optional(
|
||||
record,
|
||||
{
|
||||
"schemaVersion",
|
||||
"recordId",
|
||||
"role",
|
||||
"lifecycleStatus",
|
||||
"assetRef",
|
||||
"assetVersion",
|
||||
"artifactHash",
|
||||
"evidenceRefs",
|
||||
},
|
||||
_RECORD_KEYS - {
|
||||
"schemaVersion",
|
||||
"recordId",
|
||||
"role",
|
||||
"lifecycleStatus",
|
||||
"assetRef",
|
||||
"assetVersion",
|
||||
"artifactHash",
|
||||
"evidenceRefs",
|
||||
},
|
||||
record_id=record_id,
|
||||
)
|
||||
current_id = record.get("recordId")
|
||||
if record.get("schemaVersion") != "ReferenceAssetRecord/2" or not _is_identifier(current_id):
|
||||
# 只有已经通过 identifier 校验的上下文 ID 才能进入异常,避免回显原始输入。
|
||||
safe_record_id = current_id if _is_identifier(current_id) else (
|
||||
record_id if _is_identifier(record_id) else None
|
||||
)
|
||||
_fail("reference_registry_invalid", safe_record_id)
|
||||
if record.get("role") not in _ALLOWED_ROLES or record.get("lifecycleStatus") not in _ALLOWED_LIFECYCLE:
|
||||
_fail("reference_registry_invalid", current_id)
|
||||
_path(record.get("assetRef"), record_id=current_id)
|
||||
if not isinstance(record.get("assetVersion"), str) or not record["assetVersion"]:
|
||||
_fail("reference_registry_invalid", current_id)
|
||||
if not _is_sha256(record.get("artifactHash")):
|
||||
_fail("reference_registry_invalid", current_id)
|
||||
for field in ("consumerRef", "signedBy"):
|
||||
if not _validate_nullable_string(record.get(field)):
|
||||
_fail("reference_registry_invalid", current_id)
|
||||
signed_at = record.get("signedAt")
|
||||
if signed_at is not None and (not isinstance(signed_at, str) or _SIGNED_AT_RE.fullmatch(signed_at) is None):
|
||||
_fail("reference_registry_invalid", current_id)
|
||||
design_ref = record.get("designRef")
|
||||
if design_ref is not None:
|
||||
if not isinstance(design_ref, list) or not design_ref or not all(
|
||||
isinstance(value, str) and value for value in design_ref
|
||||
):
|
||||
_fail("reference_registry_invalid", current_id)
|
||||
for value in design_ref:
|
||||
_path(value, record_id=current_id)
|
||||
evidence_refs = record.get("evidenceRefs")
|
||||
if not isinstance(evidence_refs, list) or not all(isinstance(value, str) and value for value in evidence_refs):
|
||||
_fail("reference_registry_invalid", current_id)
|
||||
for value in evidence_refs:
|
||||
_path(value, record_id=current_id)
|
||||
for field in ("artifactRef", "consumptionManifestRef"):
|
||||
value = record.get(field)
|
||||
if value is not None:
|
||||
_path(value, record_id=current_id)
|
||||
if not (record.get("consumptionManifestHash") is None or _is_sha256(record.get("consumptionManifestHash"))):
|
||||
_fail("reference_registry_invalid", current_id)
|
||||
|
||||
active = record["lifecycleStatus"] == "active"
|
||||
historical = record["lifecycleStatus"] == "retired" and any(
|
||||
record.get(field) is not None for field in ("consumerRef", "signedBy", "signedAt")
|
||||
)
|
||||
if active or historical:
|
||||
required_identity = (
|
||||
"consumerRef",
|
||||
"signedBy",
|
||||
"signedAt",
|
||||
"artifactRef",
|
||||
"consumptionManifestRef",
|
||||
"consumptionManifestHash",
|
||||
)
|
||||
if any(record.get(field) in (None, "") for field in required_identity):
|
||||
_fail("reference_registry_invalid", current_id)
|
||||
if active and record["role"] == "game_content_gold":
|
||||
if not isinstance(design_ref, list) or not design_ref:
|
||||
_fail("reference_registry_invalid", current_id)
|
||||
|
||||
|
||||
def _validate_registry(registry: dict) -> dict[str, dict]:
|
||||
"""校验 Registry/2 的 strict 结构、recordId 唯一性和迁移语义。"""
|
||||
_keys_optional(
|
||||
registry,
|
||||
{"schemaVersion", "registryVersion", "sourceOfTruth", "records"},
|
||||
{"migrationNotes"},
|
||||
)
|
||||
if registry.get("schemaVersion") != "ReferenceAssetRegistry/2":
|
||||
_fail("reference_registry_invalid")
|
||||
if not isinstance(registry.get("registryVersion"), str) or not registry["registryVersion"]:
|
||||
_fail("reference_registry_invalid")
|
||||
if not isinstance(registry.get("sourceOfTruth"), str) or not registry["sourceOfTruth"]:
|
||||
_fail("reference_registry_invalid")
|
||||
records = registry.get("records")
|
||||
if not isinstance(records, list):
|
||||
_fail("reference_registry_invalid")
|
||||
notes = registry.get("migrationNotes", {})
|
||||
if not isinstance(notes, dict) or any(
|
||||
not _is_identifier(key) or not isinstance(value, str) or not value for key, value in notes.items()
|
||||
):
|
||||
_fail("reference_registry_invalid")
|
||||
by_id: dict[str, dict] = {}
|
||||
for record in records:
|
||||
_validate_record(record)
|
||||
current_id = record["recordId"]
|
||||
if current_id in by_id:
|
||||
_fail("reference_registry_invalid", current_id)
|
||||
by_id[current_id] = record
|
||||
for key in notes:
|
||||
if key != "_registry" and key not in by_id:
|
||||
_fail("reference_registry_invalid", key)
|
||||
if registry["registryVersion"].endswith(".migration-list") and any(
|
||||
record.get("lifecycleStatus") == "active" for record in records
|
||||
):
|
||||
_fail("reference_registry_invalid")
|
||||
return by_id
|
||||
|
||||
|
||||
def _validate_policy(policy: dict, requested_policy_id: str, requested_mode: str) -> None:
|
||||
"""校验 release 绑定 policy;当前冻结策略另执行固定身份闭包。"""
|
||||
_keys(policy, _POLICY_KEYS, code="reference_policy_missing")
|
||||
if policy.get("schemaVersion") != "ReferenceAssetConsumptionPolicy/1":
|
||||
_fail("reference_policy_missing")
|
||||
if policy.get("policyId") != requested_policy_id or not _is_identifier(policy.get("policyId")):
|
||||
_fail("reference_policy_missing")
|
||||
if not _is_identifier(policy.get("recordId")) or policy.get("role") not in _ALLOWED_ROLES:
|
||||
_fail("reference_policy_missing")
|
||||
if not isinstance(policy.get("consumerRef"), str) or not policy["consumerRef"]:
|
||||
_fail("reference_policy_missing")
|
||||
if not isinstance(policy.get("route"), str) or not policy["route"] or policy.get("autoSelect") is not False:
|
||||
_fail("reference_policy_missing")
|
||||
if policy.get("mode") != POLICY_MODE:
|
||||
_fail("reference_declaration_mismatch", policy["recordId"])
|
||||
if requested_policy_id == POLICY_ID:
|
||||
fixed = {
|
||||
"recordId": POLICY_RECORD_ID,
|
||||
"role": POLICY_ROLE,
|
||||
"consumerRef": POLICY_CONSUMER_REF,
|
||||
"route": "survivor-gold",
|
||||
}
|
||||
if any(policy.get(field) != expected for field, expected in fixed.items()):
|
||||
_fail("reference_policy_missing")
|
||||
if requested_mode != policy["mode"]:
|
||||
_fail("reference_declaration_mismatch", policy["recordId"])
|
||||
|
||||
|
||||
def _declarations_match(declarations, record: dict, policy: dict) -> bool:
|
||||
"""把调用方声明投影成 recordId/role/consumerRef 三元组后逐项对账。"""
|
||||
if declarations is None:
|
||||
return True
|
||||
if isinstance(declarations, Mapping):
|
||||
record_ids = declarations.get("referenceAssetRecordIds", declarations.get("recordIds"))
|
||||
if record_ids is None and "recordId" in declarations:
|
||||
record_ids = [declarations.get("recordId")]
|
||||
consumer_ref = declarations.get("consumerRef")
|
||||
role = declarations.get("role")
|
||||
elif isinstance(declarations, Sequence) and not isinstance(declarations, (str, bytes, bytearray)):
|
||||
record_ids = list(declarations)
|
||||
consumer_ref = None
|
||||
role = None
|
||||
else:
|
||||
return False
|
||||
if record_ids != [record["recordId"]] or consumer_ref not in (None, policy["consumerRef"]):
|
||||
return False
|
||||
return role in (None, policy["role"])
|
||||
|
||||
|
||||
def _under(path: str, root: str) -> bool:
|
||||
"""判断仓内路径是否位于目录根或等于批准的文件引用。"""
|
||||
return path == root or path.startswith(root + "/")
|
||||
|
||||
|
||||
def _validate_manifest(manifest: dict, raw: bytes, *, record_id: str, allowed_root: str,
|
||||
design_refs: Sequence[str]) -> list[dict]:
|
||||
"""校验 manifest canonical 字节、排序、路径范围和单记录资源预算。"""
|
||||
if len(raw) > MAX_MANIFEST_BYTES:
|
||||
_fail("reference_oversize", record_id)
|
||||
if raw != canonical_json_bytes(manifest):
|
||||
_fail("reference_manifest_hash_mismatch", record_id)
|
||||
_keys(manifest, _MANIFEST_KEYS, record_id=record_id, code="reference_manifest_hash_mismatch")
|
||||
if manifest.get("schemaVersion") != "ReferenceAssetConsumptionManifest/1":
|
||||
_fail("reference_manifest_hash_mismatch", record_id)
|
||||
if not _is_identifier(manifest.get("manifestId")):
|
||||
_fail("reference_manifest_hash_mismatch", record_id)
|
||||
if manifest.get("canonicalization") != "reference-asset-consumption-manifest/1":
|
||||
_fail("reference_manifest_hash_mismatch", record_id)
|
||||
entries = manifest.get("entries")
|
||||
if not isinstance(entries, list) or not entries:
|
||||
_fail("reference_manifest_hash_mismatch", record_id)
|
||||
if len(entries) > MAX_MANIFEST_ENTRIES:
|
||||
_fail("reference_oversize", record_id)
|
||||
|
||||
normalized_paths: list[str] = []
|
||||
seen: set[str] = set()
|
||||
total_declared = 0
|
||||
for entry in entries:
|
||||
_keys(entry, _MANIFEST_ENTRY_KEYS, record_id=record_id, code="reference_manifest_hash_mismatch")
|
||||
path = _path(entry.get("path"), record_id=record_id)
|
||||
if path in seen:
|
||||
_fail("reference_path_invalid", record_id, path)
|
||||
seen.add(path)
|
||||
normalized_paths.append(path)
|
||||
size = entry.get("size")
|
||||
if not isinstance(size, int) or isinstance(size, bool) or size < 0 or size > 9007199254740991:
|
||||
_fail("reference_manifest_hash_mismatch", record_id, path)
|
||||
if not _is_sha256(entry.get("sha256")):
|
||||
_fail("reference_manifest_hash_mismatch", record_id, path)
|
||||
if size > MAX_FILE_BYTES:
|
||||
_fail("reference_oversize", record_id, path)
|
||||
total_declared += size
|
||||
if total_declared > MAX_RECORD_BYTES:
|
||||
_fail("reference_oversize", record_id, path)
|
||||
if not _under(path, allowed_root) and path not in design_refs:
|
||||
_fail("reference_path_escape", record_id, path)
|
||||
if normalized_paths != sorted(normalized_paths, key=lambda value: value.encode("utf-8")):
|
||||
_fail("reference_path_invalid", record_id)
|
||||
return entries
|
||||
|
||||
|
||||
def _stage_one_policy(
|
||||
policy_id: str,
|
||||
*,
|
||||
release_ref: str,
|
||||
expected_release_hash: str,
|
||||
trusted_root,
|
||||
mode: str,
|
||||
declarations=None,
|
||||
) -> _StagedReferenceAssets:
|
||||
"""在已锚定根 fd 上完成单条 policy 验证,仅返回 batch 内部暂存值。"""
|
||||
if not _is_sha256(expected_release_hash):
|
||||
_fail("reference_registry_untrusted")
|
||||
|
||||
release_ref = _path(release_ref)
|
||||
root_fd = _open_trusted_root(trusted_root)
|
||||
try:
|
||||
release, _, release_hash = _read_json_file(
|
||||
root_fd,
|
||||
release_ref,
|
||||
record_id=None,
|
||||
limits={"max_files": 1, "max_file_bytes": MAX_MANIFEST_BYTES, "max_record_bytes": MAX_MANIFEST_BYTES},
|
||||
)
|
||||
if release_hash != expected_release_hash:
|
||||
_fail("reference_registry_untrusted", logical_path=release_ref)
|
||||
_validate_release(release)
|
||||
|
||||
registry_ref = _path(release["registryRef"])
|
||||
registry, _, registry_hash = _read_json_file(
|
||||
root_fd,
|
||||
registry_ref,
|
||||
record_id=None,
|
||||
limits={"max_files": 1, "max_file_bytes": MAX_RECORD_BYTES, "max_record_bytes": MAX_RECORD_BYTES},
|
||||
)
|
||||
if registry_hash != release["registryHash"]:
|
||||
_fail("reference_registry_untrusted", logical_path=registry_ref)
|
||||
by_id = _validate_registry(registry)
|
||||
|
||||
policy_ref = _path(release["policyRef"])
|
||||
try:
|
||||
policy, _, policy_hash = _read_json_file(
|
||||
root_fd,
|
||||
policy_ref,
|
||||
record_id=None,
|
||||
limits={"max_files": 1, "max_file_bytes": MAX_MANIFEST_BYTES, "max_record_bytes": MAX_MANIFEST_BYTES},
|
||||
missing_code="reference_policy_missing",
|
||||
)
|
||||
except ReferenceAssetGateError as exc:
|
||||
if exc.code == "reference_registry_invalid":
|
||||
_fail("reference_policy_missing", logical_path=policy_ref)
|
||||
raise
|
||||
if policy_hash != release["policyHash"]:
|
||||
_fail("reference_registry_untrusted", logical_path=policy_ref)
|
||||
_validate_policy(policy, policy_id, mode)
|
||||
|
||||
record = by_id.get(policy["recordId"])
|
||||
if record is None:
|
||||
_fail("reference_registry_invalid", policy["recordId"])
|
||||
if record.get("lifecycleStatus") != "active":
|
||||
_fail("reference_registry_invalid", record["recordId"])
|
||||
if record.get("role") != policy["role"] or record.get("consumerRef") != policy["consumerRef"]:
|
||||
_fail("reference_declaration_mismatch", record["recordId"])
|
||||
if not _declarations_match(declarations, record, policy):
|
||||
_fail("reference_declaration_mismatch", record["recordId"])
|
||||
|
||||
record_id = record["recordId"]
|
||||
asset_root = _path(record["assetRef"], record_id=record_id)
|
||||
artifact_ref = _path(record.get("artifactRef"), record_id=record_id)
|
||||
manifest_ref = _path(record.get("consumptionManifestRef"), record_id=record_id)
|
||||
if not _under(artifact_ref, asset_root) or not _under(manifest_ref, asset_root):
|
||||
_fail("reference_path_escape", record_id)
|
||||
design_refs = tuple(record.get("designRef") or ())
|
||||
|
||||
# 先读取并核验 bundle;manifest 使用独立的 1 MiB 读前预算,避免把超限清单
|
||||
# 先读入内存。两次读取仍共用同一个 trusted root fd 和逐级 O_NOFOLLOW 边界。
|
||||
artifact_files = _read_files(
|
||||
root_fd,
|
||||
[artifact_ref],
|
||||
limits={
|
||||
"max_files": 1,
|
||||
"max_file_bytes": MAX_FILE_BYTES,
|
||||
"max_record_bytes": MAX_RECORD_BYTES,
|
||||
},
|
||||
record_id=record_id,
|
||||
)
|
||||
observed_artifact_hash = hashlib.sha256(artifact_files[artifact_ref]).hexdigest()
|
||||
if observed_artifact_hash != record["artifactHash"]:
|
||||
_fail("reference_artifact_hash_mismatch", record_id, artifact_ref)
|
||||
manifest_files = _read_files(
|
||||
root_fd,
|
||||
[manifest_ref],
|
||||
limits={
|
||||
"max_files": 1,
|
||||
"max_file_bytes": MAX_MANIFEST_BYTES,
|
||||
"max_record_bytes": MAX_MANIFEST_BYTES,
|
||||
},
|
||||
record_id=record_id,
|
||||
)
|
||||
manifest_raw = manifest_files[manifest_ref]
|
||||
observed_manifest_hash = hashlib.sha256(manifest_raw).hexdigest()
|
||||
if observed_manifest_hash != record["consumptionManifestHash"]:
|
||||
_fail("reference_manifest_hash_mismatch", record_id, manifest_ref)
|
||||
manifest = _load_json(
|
||||
manifest_raw,
|
||||
code="reference_manifest_hash_mismatch",
|
||||
record_id=record_id,
|
||||
logical_path=manifest_ref,
|
||||
)
|
||||
entries = _validate_manifest(
|
||||
manifest,
|
||||
manifest_raw,
|
||||
record_id=record_id,
|
||||
allowed_root=asset_root,
|
||||
design_refs=design_refs,
|
||||
)
|
||||
entry_paths = [entry["path"] for entry in entries]
|
||||
selected = _read_files(
|
||||
root_fd,
|
||||
entry_paths,
|
||||
limits={
|
||||
"max_files": MAX_MANIFEST_ENTRIES,
|
||||
"max_file_bytes": MAX_FILE_BYTES,
|
||||
"max_record_bytes": MAX_RECORD_BYTES,
|
||||
},
|
||||
record_id=record_id,
|
||||
)
|
||||
for entry in entries:
|
||||
path = entry["path"]
|
||||
content = selected[path]
|
||||
if len(content) != entry["size"] or hashlib.sha256(content).hexdigest() != entry["sha256"]:
|
||||
_fail("reference_entry_hash_mismatch", record_id, path)
|
||||
|
||||
immutable_files = MappingProxyType(dict(selected))
|
||||
receipt = {
|
||||
"schemaVersion": "ReferenceAssetVerificationReceipt/1",
|
||||
"receiptId": f"receipt-{policy_id}-{record_id}-{observed_manifest_hash[:12]}",
|
||||
"releaseRef": release_ref,
|
||||
"registryVersion": registry["registryVersion"],
|
||||
"expectedRegistryHash": release["registryHash"],
|
||||
"observedRegistryHash": registry_hash,
|
||||
"policyHash": policy_hash,
|
||||
"verifierVersion": release["verifierVersion"],
|
||||
"trustedRootId": release["trustedRootId"],
|
||||
"recordId": record_id,
|
||||
"role": record["role"],
|
||||
"consumerRef": record["consumerRef"],
|
||||
"artifactRef": artifact_ref,
|
||||
"consumptionManifestRef": manifest_ref,
|
||||
"expected": {
|
||||
"artifactHash": record["artifactHash"],
|
||||
"consumptionManifestHash": record["consumptionManifestHash"],
|
||||
},
|
||||
"observed": {
|
||||
"artifactHash": observed_artifact_hash,
|
||||
"consumptionManifestHash": observed_manifest_hash,
|
||||
},
|
||||
"finalSnapshotHash": snapshot_hash(immutable_files),
|
||||
}
|
||||
return _StagedReferenceAssets(
|
||||
record=dict(record),
|
||||
files=immutable_files,
|
||||
receipt=receipt,
|
||||
roots=tuple((asset_root, *design_refs)),
|
||||
)
|
||||
finally:
|
||||
os.close(root_fd)
|
||||
|
||||
|
||||
def verify_policy(
|
||||
policy_id: str,
|
||||
*,
|
||||
release_ref: str,
|
||||
expected_release_hash: str,
|
||||
trusted_root,
|
||||
mode: str,
|
||||
declarations=None,
|
||||
) -> VerifiedReferenceAssets:
|
||||
"""验证当前唯一受批准 policy,并委托单元素原子 batch。"""
|
||||
if policy_id != POLICY_ID:
|
||||
_fail("reference_policy_missing")
|
||||
if mode != POLICY_MODE:
|
||||
_fail("reference_declaration_mismatch")
|
||||
return verify_policies([{
|
||||
"policy_id": policy_id,
|
||||
"release_ref": release_ref,
|
||||
"expected_release_hash": expected_release_hash,
|
||||
"trusted_root": trusted_root,
|
||||
"mode": mode,
|
||||
"declarations": declarations,
|
||||
}])
|
||||
|
||||
|
||||
def verify_policies(requests: Sequence[Mapping[str, object]]) -> VerifiedReferenceAssets:
|
||||
"""原子验证并合并多条 policy,任一失败都不会构造公开结果。"""
|
||||
if not isinstance(requests, Sequence) or isinstance(requests, (str, bytes, bytearray)) or not requests:
|
||||
_fail("reference_policy_missing")
|
||||
|
||||
required_keys = {
|
||||
"policy_id",
|
||||
"release_ref",
|
||||
"expected_release_hash",
|
||||
"trusted_root",
|
||||
"mode",
|
||||
}
|
||||
allowed_keys = required_keys | {"declarations"}
|
||||
merged_files: dict[str, bytes] = {}
|
||||
merged_records: list[Mapping[str, object]] = []
|
||||
merged_receipts: list[Mapping[str, object]] = []
|
||||
merged_roots: dict[str, tuple[str, ...]] = {}
|
||||
total_bytes = 0
|
||||
|
||||
for request in requests:
|
||||
if not isinstance(request, Mapping) or not required_keys.issubset(request) or set(request) - allowed_keys:
|
||||
_fail("reference_declaration_mismatch")
|
||||
item = _stage_one_policy(
|
||||
request["policy_id"],
|
||||
release_ref=request["release_ref"],
|
||||
expected_release_hash=request["expected_release_hash"],
|
||||
trusted_root=request["trusted_root"],
|
||||
mode=request["mode"],
|
||||
declarations=request.get("declarations"),
|
||||
)
|
||||
record_id = item.record["recordId"]
|
||||
if record_id in merged_roots:
|
||||
_fail("reference_registry_invalid", record_id)
|
||||
for path, content in item.files.items():
|
||||
existing = merged_files.get(path)
|
||||
if existing is not None:
|
||||
if existing != content:
|
||||
_fail("reference_entry_hash_mismatch", record_id, path)
|
||||
continue
|
||||
total_bytes += len(content)
|
||||
if total_bytes > MAX_TOTAL_BYTES:
|
||||
_fail("reference_oversize", record_id, path)
|
||||
merged_files[path] = content
|
||||
# 记录和回执在当前 stage 合并成功后立即冻结,避免积存可变暂存对象。
|
||||
merged_records.append(_freeze(item.record))
|
||||
merged_receipts.append(_freeze(item.receipt))
|
||||
merged_roots[record_id] = tuple(item.roots)
|
||||
|
||||
immutable_files = MappingProxyType(dict(merged_files))
|
||||
return VerifiedReferenceAssets(
|
||||
constraint_records=tuple(merged_records),
|
||||
files=immutable_files,
|
||||
receipts=tuple(merged_receipts),
|
||||
snapshot_hash=snapshot_hash(immutable_files),
|
||||
reference_roots=_freeze(merged_roots),
|
||||
)
|
||||
|
||||
|
||||
__all__ = [
|
||||
"MAX_FILE_BYTES",
|
||||
"MAX_MANIFEST_BYTES",
|
||||
"MAX_MANIFEST_ENTRIES",
|
||||
"MAX_RECORD_BYTES",
|
||||
"MAX_TOTAL_BYTES",
|
||||
"POLICY_ID",
|
||||
"ReferenceAssetError",
|
||||
"ReferenceAssetGateError",
|
||||
"ReferenceAssetVerificationError",
|
||||
"VerifiedReferenceAssets",
|
||||
"canonical_json_bytes",
|
||||
"canonical_snapshot_hash",
|
||||
"snapshot_hash",
|
||||
"to_json_value",
|
||||
"verify_policy",
|
||||
"verify_policies",
|
||||
]
|
||||
2112
cheap-worker/tests/test_acceptance_v3.py
Normal file
2112
cheap-worker/tests/test_acceptance_v3.py
Normal file
File diff suppressed because it is too large
Load Diff
463
cheap-worker/tests/test_baseline_gates.py
Normal file
463
cheap-worker/tests/test_baseline_gates.py
Normal file
@ -0,0 +1,463 @@
|
||||
"""三批基线闸门测试:fresh25 阈值边界 + historical11 固定预期表 + shadow20 达标断言(确定性,零模型)。
|
||||
|
||||
阈值严格按设计档 §5.3/§5.2/§5.4,本测试固化边界(19/25 拒 vs 20/25 过、品类 2/5 拒、成本 ¥1.5/¥15 临界、
|
||||
假阳放行 0、needs_human 不自动 accept、tester_error=0、inconclusive≤1、固定六项零漂移、fresh25 硬前置)。
|
||||
"""
|
||||
|
||||
import asyncio
|
||||
import json
|
||||
import sys
|
||||
from pathlib import Path
|
||||
|
||||
import pytest
|
||||
|
||||
sys.path.insert(0, str(Path(__file__).resolve().parents[1]))
|
||||
|
||||
from baseline_gates import ( # noqa: E402
|
||||
FRESH25_GENRES, build_shadow_plan, evaluate_fresh25_gate, evaluate_historical_gate,
|
||||
evaluate_shadow20_gate, load_historical_expectations, run_shadow_batch,
|
||||
)
|
||||
|
||||
|
||||
# ────────────────────────── fresh25 ──────────────────────────
|
||||
|
||||
def _fresh_rows(accepted_per_genre=(4, 4, 4, 4, 4), *, first_pass=18, repair=0, rescued=0,
|
||||
run_cost=1.0, chain_cost=5.0, drop=0, dup_gid=False):
|
||||
"""按品类构造 25 局批结果行(gid/genre 与 hard_genre_batch 同 schema)。"""
|
||||
rows = []
|
||||
idx = 0
|
||||
for gi, genre in enumerate(FRESH25_GENRES):
|
||||
for rnd in range(1, 6):
|
||||
rows.append({
|
||||
"gid": f"{genre}-r{rnd}", "genre": genre, "round": rnd,
|
||||
"accepted": rnd <= accepted_per_genre[gi],
|
||||
"firstPassAccepted": idx < first_pass,
|
||||
"repairAttempted": idx < repair,
|
||||
"rescuedByRoll": 2 if idx < rescued else None,
|
||||
"acceptedAfterRepair": False,
|
||||
"outcome": "accept" if rnd <= accepted_per_genre[gi] else "reject",
|
||||
"acceptanceCostRmb": run_cost, "parentChainCostRmb": chain_cost,
|
||||
})
|
||||
idx += 1
|
||||
if drop:
|
||||
rows = rows[:-drop]
|
||||
if dup_gid:
|
||||
rows[-1]["gid"] = rows[0]["gid"]
|
||||
return rows
|
||||
|
||||
|
||||
def test_fresh25_all_green_passes():
|
||||
result = evaluate_fresh25_gate(_fresh_rows())
|
||||
assert result["pass"] is True
|
||||
assert result["metrics"]["accepted"] == 20
|
||||
assert all(c["pass"] for c in result["checks"])
|
||||
|
||||
|
||||
def test_fresh25_boundary_19_of_25_rejected():
|
||||
"""accepted 19/25(4,4,4,4,3)→ 总数闸拒;品类 3/5 仍达标,只挂总数一项。"""
|
||||
result = evaluate_fresh25_gate(_fresh_rows(accepted_per_genre=(4, 4, 4, 4, 3)))
|
||||
assert result["pass"] is False
|
||||
by_name = {c["name"]: c for c in result["checks"]}
|
||||
assert by_name["acceptedTotal"]["pass"] is False
|
||||
assert by_name["acceptedTotal"]["actual"] == "19/25"
|
||||
assert by_name["perGenreMin"]["pass"] is True
|
||||
|
||||
|
||||
def test_fresh25_boundary_20_of_25_passes():
|
||||
result = evaluate_fresh25_gate(_fresh_rows(accepted_per_genre=(4, 4, 4, 4, 4)))
|
||||
by_name = {c["name"]: c for c in result["checks"]}
|
||||
assert by_name["acceptedTotal"]["pass"] is True
|
||||
assert result["pass"] is True
|
||||
|
||||
|
||||
def test_fresh25_genre_2_of_5_rejected_even_with_20_total():
|
||||
"""总数 20/25 达标但某品类 2/5 → 品类闸拒(任一品类不低于 3/5)。"""
|
||||
result = evaluate_fresh25_gate(_fresh_rows(accepted_per_genre=(5, 5, 5, 3, 2)))
|
||||
assert result["pass"] is False
|
||||
by_name = {c["name"]: c for c in result["checks"]}
|
||||
assert by_name["acceptedTotal"]["pass"] is True # 总数 20 达标
|
||||
assert by_name["perGenreMin"]["pass"] is False
|
||||
assert "sim-business" in by_name["perGenreMin"]["detail"]
|
||||
|
||||
|
||||
def test_fresh25_first_pass_boundary():
|
||||
"""firstPassAccepted 18 过 / 17 拒。"""
|
||||
assert evaluate_fresh25_gate(_fresh_rows(first_pass=18))["pass"] is True
|
||||
result = evaluate_fresh25_gate(_fresh_rows(first_pass=17))
|
||||
assert result["pass"] is False
|
||||
assert {c["name"]: c for c in result["checks"]}["firstPassAccepted"]["actual"] == "17/25"
|
||||
|
||||
|
||||
def test_fresh25_repair_rate_boundary():
|
||||
"""writer repair 启动 5 过 / 6 拒。"""
|
||||
assert evaluate_fresh25_gate(_fresh_rows(repair=5))["pass"] is True
|
||||
assert evaluate_fresh25_gate(_fresh_rows(repair=6))["pass"] is False
|
||||
|
||||
|
||||
def test_fresh25_rescued_boundary():
|
||||
"""rescuedByRoll 5 过 / 6 拒。"""
|
||||
assert evaluate_fresh25_gate(_fresh_rows(rescued=5))["pass"] is True
|
||||
assert evaluate_fresh25_gate(_fresh_rows(rescued=6))["pass"] is False
|
||||
|
||||
|
||||
def test_fresh25_run_cost_boundary():
|
||||
"""单局验收成本 ¥1.5 过 / ¥1.5001 拒。"""
|
||||
assert evaluate_fresh25_gate(_fresh_rows(run_cost=1.5))["pass"] is True
|
||||
result = evaluate_fresh25_gate(_fresh_rows(run_cost=1.5001))
|
||||
assert result["pass"] is False
|
||||
assert {c["name"]: c for c in result["checks"]}["perRunCost"]["pass"] is False
|
||||
|
||||
|
||||
def test_fresh25_chain_cost_boundary():
|
||||
"""parentRun 全链 ¥15 过 / ¥15.0001 拒。"""
|
||||
assert evaluate_fresh25_gate(_fresh_rows(chain_cost=15.0))["pass"] is True
|
||||
assert evaluate_fresh25_gate(_fresh_rows(chain_cost=15.0001))["pass"] is False
|
||||
|
||||
|
||||
def test_fresh25_missing_cost_fails_closed():
|
||||
"""缺验收成本 = 无法证明达标 → fail-closed 拒。"""
|
||||
rows = _fresh_rows()
|
||||
rows[3]["acceptanceCostRmb"] = None
|
||||
result = evaluate_fresh25_gate(rows)
|
||||
assert result["pass"] is False
|
||||
assert {c["name"]: c for c in result["checks"]}["perRunCost"]["pass"] is False
|
||||
|
||||
|
||||
def test_fresh25_nan_cost_treated_as_missing():
|
||||
rows = _fresh_rows()
|
||||
rows[0]["parentChainCostRmb"] = float("nan")
|
||||
assert evaluate_fresh25_gate(rows)["pass"] is False
|
||||
|
||||
|
||||
def test_fresh25_incomplete_sample_rejected():
|
||||
"""24 局(分母不完整)→ 样本闸拒,不能拿残缺基线冒充达标。"""
|
||||
result = evaluate_fresh25_gate(_fresh_rows(drop=1))
|
||||
assert result["pass"] is False
|
||||
assert {c["name"]: c for c in result["checks"]}["sampleSize"]["pass"] is False
|
||||
|
||||
|
||||
def test_fresh25_duplicate_gid_rejected():
|
||||
"""同 gid 重跑洗数字 → 样本闸拒(distinct gid 纪律)。"""
|
||||
result = evaluate_fresh25_gate(_fresh_rows(dup_gid=True))
|
||||
assert result["pass"] is False
|
||||
assert {c["name"]: c for c in result["checks"]}["sampleSize"]["pass"] is False
|
||||
|
||||
|
||||
def test_fresh25_accepted_after_repair_is_observation_not_numerator():
|
||||
"""acceptedAfterRepair 只作修复救回率观测,绝不当成功率分子。"""
|
||||
result = evaluate_fresh25_gate(_fresh_rows())
|
||||
assert result["metrics"]["acceptedAfterRepair_observation"] == 0
|
||||
assert "acceptedAfterRepair" not in {c["name"] for c in result["checks"]} # 不作闸门分子
|
||||
|
||||
|
||||
# ────────────────────────── historical11 ──────────────────────────
|
||||
|
||||
def _expected_matched_mapping():
|
||||
"""与固定预期表逐局相符的重放结果(needs_human 局给 inconclusive,仍挂起待定标)。"""
|
||||
return {
|
||||
"narrative-r1": "accept", "narrative-r2": "accept", "narrative-r3": "accept",
|
||||
"trpg-r1": "reject", "puzzle-r1": "inconclusive", "trpg-r2": "reject",
|
||||
"heritage-r2": "inconclusive", "sim-business-r2": "reject",
|
||||
"puzzle-r2": "reject",
|
||||
"heritage-r1": "inconclusive", "sim-business-r1": "inconclusive",
|
||||
}
|
||||
|
||||
|
||||
def _rows_from(mapping):
|
||||
return [{"gid": gid, "outcome": outcome} for gid, outcome in mapping.items()]
|
||||
|
||||
|
||||
def test_expectations_fixture_is_eleven_well_formed_games():
|
||||
data = load_historical_expectations()
|
||||
rows = data["expectations"]
|
||||
assert len(rows) == 11
|
||||
assert {r["expected"] for r in rows} <= {"accept", "reject", "not_accept", "needs_human"}
|
||||
categories = {r["category"]: sum(1 for x in rows if x["category"] == r["category"]) for r in rows}
|
||||
assert categories["positive_control"] == 3 # 3 narrative 正例
|
||||
assert categories["known_false_positive"] == 5 # 5 旧假阳
|
||||
assert categories["true_bug"] == 1 # puzzle-r2
|
||||
assert categories["suspected_false_negative"] == 2 # 2 疑似假阴
|
||||
|
||||
|
||||
def test_historical_matched_but_blocked_on_human_calibration():
|
||||
"""逐局相符 + needs_human 未定标 → 闸门不自动 PASS,挂起 2 局交创始人(红线)。"""
|
||||
result = evaluate_historical_gate(_rows_from(_expected_matched_mapping()))
|
||||
assert result["pass"] is False
|
||||
assert result["falsePositiveReleased"] == []
|
||||
assert sorted(result["blockedOnHumanCalibration"]) == ["heritage-r1", "sim-business-r1"]
|
||||
by_name = {c["name"]: c for c in result["checks"]}
|
||||
assert by_name["falsePositiveRelease"]["pass"] is True
|
||||
assert by_name["perGameExpectation"]["pass"] is True
|
||||
assert by_name["humanCalibrationSettled"]["pass"] is False
|
||||
|
||||
|
||||
def test_historical_known_false_positive_accept_is_release_fail():
|
||||
"""旧假阳无新硬证被 accept = 确认假阳放行,闸门红线。"""
|
||||
mapping = _expected_matched_mapping()
|
||||
mapping["sim-business-r2"] = "accept"
|
||||
result = evaluate_historical_gate(_rows_from(mapping))
|
||||
assert result["pass"] is False
|
||||
assert result["falsePositiveReleased"] == ["sim-business-r2"]
|
||||
assert {c["name"]: c for c in result["checks"]}["falsePositiveRelease"]["pass"] is False
|
||||
|
||||
|
||||
def test_historical_true_bug_accept_is_release_fail():
|
||||
"""puzzle-r2 真 bug 被 accept = 假阳放行。"""
|
||||
mapping = _expected_matched_mapping()
|
||||
mapping["puzzle-r2"] = "accept"
|
||||
result = evaluate_historical_gate(_rows_from(mapping))
|
||||
assert "puzzle-r2" in result["falsePositiveReleased"]
|
||||
|
||||
|
||||
def test_historical_true_bug_inconclusive_is_mismatch():
|
||||
"""puzzle-r2 固定预期 reject,inconclusive 与预期不符(查验收器为何拿不到硬证)。"""
|
||||
mapping = _expected_matched_mapping()
|
||||
mapping["puzzle-r2"] = "inconclusive"
|
||||
result = evaluate_historical_gate(_rows_from(mapping))
|
||||
assert result["pass"] is False
|
||||
by_name = {c["name"]: c for c in result["checks"]}
|
||||
assert by_name["perGameExpectation"]["pass"] is False
|
||||
assert "puzzle-r2" in by_name["perGameExpectation"]["detail"]
|
||||
|
||||
|
||||
def test_historical_positive_control_must_still_accept():
|
||||
"""narrative 正例应保留:v3 判 inconclusive = 验收器回退,逐局预期不符。"""
|
||||
mapping = _expected_matched_mapping()
|
||||
mapping["narrative-r1"] = "inconclusive"
|
||||
result = evaluate_historical_gate(_rows_from(mapping))
|
||||
assert result["pass"] is False
|
||||
assert "narrative-r1" in {c["name"]: c for c in result["checks"]}["perGameExpectation"]["detail"]
|
||||
|
||||
|
||||
def test_historical_needs_human_accept_is_pre_calibration_fail():
|
||||
"""疑似假阴未定标即 accept → 定标前不得自动接受(红线)。"""
|
||||
mapping = _expected_matched_mapping()
|
||||
mapping["heritage-r1"] = "accept"
|
||||
result = evaluate_historical_gate(_rows_from(mapping))
|
||||
assert result["pass"] is False
|
||||
assert {c["name"]: c for c in result["checks"]}["preCalibrationAccept"]["pass"] is False
|
||||
|
||||
|
||||
def test_historical_tester_error_is_mismatch_not_evidence():
|
||||
"""tester_error 不构成证据 → 逐局不符,需重跑。"""
|
||||
mapping = _expected_matched_mapping()
|
||||
mapping["trpg-r1"] = "tester_error"
|
||||
result = evaluate_historical_gate(_rows_from(mapping))
|
||||
assert result["pass"] is False
|
||||
assert {c["name"]: c for c in result["checks"]}["perGameExpectation"]["pass"] is False
|
||||
|
||||
|
||||
def test_historical_missing_game_is_coverage_fail():
|
||||
mapping = _expected_matched_mapping()
|
||||
del mapping["puzzle-r1"]
|
||||
result = evaluate_historical_gate(_rows_from(mapping))
|
||||
assert result["pass"] is False
|
||||
assert {c["name"]: c for c in result["checks"]}["coverage"]["pass"] is False
|
||||
|
||||
|
||||
def test_historical_fully_calibrated_can_auto_pass():
|
||||
"""创始人定标后(needs_human 改为固定预期)+ 逐局相符 → 闸门可自动 PASS。"""
|
||||
data = load_historical_expectations()
|
||||
calibrated = json.loads(json.dumps(data))
|
||||
for row in calibrated["expectations"]:
|
||||
if row["expected"] == "needs_human":
|
||||
row["expected"] = "accept" # 真人定标结论:其实是好游戏
|
||||
row["category"] = "calibrated"
|
||||
mapping = _expected_matched_mapping()
|
||||
mapping["heritage-r1"] = "accept"
|
||||
mapping["sim-business-r1"] = "accept"
|
||||
result = evaluate_historical_gate(_rows_from(mapping), expectations=calibrated,
|
||||
fresh25_gate={"pass": True})
|
||||
assert result["pass"] is True
|
||||
assert result["blockedOnHumanCalibration"] == []
|
||||
|
||||
|
||||
def test_historical_chain_blocked_by_failed_fresh25():
|
||||
"""链式校验:fresh25 未过 → warn + PASS 受阻(顺序建议,指标仍评估)。"""
|
||||
data = load_historical_expectations()
|
||||
calibrated = json.loads(json.dumps(data))
|
||||
for row in calibrated["expectations"]:
|
||||
if row["expected"] == "needs_human":
|
||||
row["expected"] = "reject"
|
||||
mapping = _expected_matched_mapping()
|
||||
mapping["heritage-r1"] = "reject"
|
||||
mapping["sim-business-r1"] = "reject"
|
||||
result = evaluate_historical_gate(_rows_from(mapping), expectations=calibrated,
|
||||
fresh25_gate={"pass": False})
|
||||
assert result["pass"] is False
|
||||
assert {c["name"]: c for c in result["checks"]}["chainFresh25"]["pass"] is False
|
||||
assert any("fresh25" in w for w in result["warnings"])
|
||||
|
||||
|
||||
# ────────────────────────── shadow20 ──────────────────────────
|
||||
|
||||
_IDENTITY = {"commitHash": "c0ffee00", "chromeVersion": "126.0.6478.183",
|
||||
"actorModel": "MiniMax-M3", "judgeModel": "MiniMax-M3",
|
||||
"promptVersion": "3.0.14", "configSnapshotHash": "cfg-snap-01"}
|
||||
|
||||
|
||||
def _shadow_rows(n=20, **overrides_by_index):
|
||||
rows = []
|
||||
for i in range(n):
|
||||
row = {"gid": f"shadow20-{i + 1:02d}", "outcome": "accept",
|
||||
"proofComplete": True, "confirmedFalsePositive": False,
|
||||
"problems": [], "contradictions": [], "proofMissing": False,
|
||||
"rescued": False, "discrepancy": False, "humanReviewed": True,
|
||||
"identity": dict(_IDENTITY)}
|
||||
row.update(overrides_by_index.get(i) or {})
|
||||
rows.append(row)
|
||||
return rows
|
||||
|
||||
|
||||
_F25_PASS = {"pass": True}
|
||||
_F25_FAIL = {"pass": False}
|
||||
|
||||
|
||||
def test_shadow20_all_green_passes_with_fresh25_prereq():
|
||||
result = evaluate_shadow20_gate(_shadow_rows(), fresh25_gate=_F25_PASS)
|
||||
assert result["pass"] is True
|
||||
assert all(c["pass"] for c in result["checks"])
|
||||
|
||||
|
||||
def test_shadow20_fresh25_prereq_hard_required():
|
||||
"""fresh25 未提供 → 硬前置无法验证,PASS 阻断(§5.4 顺序硬要求)。"""
|
||||
result = evaluate_shadow20_gate(_shadow_rows(), fresh25_gate=None)
|
||||
assert result["pass"] is False
|
||||
assert {c["name"]: c for c in result["checks"]}["prerequisiteFresh25"]["pass"] is False
|
||||
result2 = evaluate_shadow20_gate(_shadow_rows(), fresh25_gate=_F25_FAIL)
|
||||
assert result2["pass"] is False
|
||||
assert {c["name"]: c for c in result2["checks"]}["prerequisiteFresh25"]["actual"] == "fresh25 未达标"
|
||||
|
||||
|
||||
def test_shadow20_nineteen_samples_rejected():
|
||||
result = evaluate_shadow20_gate(_shadow_rows(n=19), fresh25_gate=_F25_PASS)
|
||||
assert result["pass"] is False
|
||||
assert {c["name"]: c for c in result["checks"]}["sampleSize"]["pass"] is False
|
||||
|
||||
|
||||
def test_shadow20_single_tester_error_rejected():
|
||||
"""20 局口径 <5% → tester_error 必须为 0:1 个即拒。"""
|
||||
rows = _shadow_rows()
|
||||
rows[7]["outcome"] = "tester_error"
|
||||
rows[7]["humanReviewed"] = True
|
||||
result = evaluate_shadow20_gate(rows, fresh25_gate=_F25_PASS)
|
||||
assert result["pass"] is False
|
||||
check = {c["name"]: c for c in result["checks"]}["testerError"]
|
||||
assert check["pass"] is False and check["actual"] == "1/20"
|
||||
|
||||
|
||||
def test_shadow20_inconclusive_boundary_one_ok_two_rejected():
|
||||
"""<10% → inconclusive 最多 1:1 个过、2 个拒。"""
|
||||
rows = _shadow_rows()
|
||||
rows[0]["outcome"] = "inconclusive"
|
||||
assert evaluate_shadow20_gate(rows, fresh25_gate=_F25_PASS)["pass"] is True
|
||||
rows[1]["outcome"] = "inconclusive"
|
||||
result = evaluate_shadow20_gate(rows, fresh25_gate=_F25_PASS)
|
||||
assert result["pass"] is False
|
||||
assert {c["name"]: c for c in result["checks"]}["inconclusive"]["actual"] == "2/20"
|
||||
|
||||
|
||||
def test_shadow20_confirmed_false_positive_rejected():
|
||||
rows = _shadow_rows()
|
||||
rows[3]["confirmedFalsePositive"] = True
|
||||
result = evaluate_shadow20_gate(rows, fresh25_gate=_F25_PASS)
|
||||
assert result["pass"] is False
|
||||
assert {c["name"]: c for c in result["checks"]}["confirmedFalsePositives"]["detail"] == "shadow20-04"
|
||||
|
||||
|
||||
def test_shadow20_accepted_proof_incomplete_rejected():
|
||||
"""accepted proof 完整率必须 100%:一局不完整即拒。"""
|
||||
rows = _shadow_rows()
|
||||
rows[5]["proofComplete"] = False
|
||||
result = evaluate_shadow20_gate(rows, fresh25_gate=_F25_PASS)
|
||||
assert result["pass"] is False
|
||||
assert "shadow20-06" in {c["name"]: c for c in result["checks"]}["proofCompleteOnAccepted"]["detail"]
|
||||
|
||||
|
||||
def test_shadow20_accepted_with_conflict_rejected():
|
||||
"""accepted 与 problems/缺证/矛盾共存 0:三种脏状态各测一遍。"""
|
||||
for field, value in (("problems", ["结论与截图矛盾"]), ("contradictions", ["事件序倒序"]),
|
||||
("proofMissing", True)):
|
||||
rows = _shadow_rows()
|
||||
rows[2][field] = value
|
||||
result = evaluate_shadow20_gate(rows, fresh25_gate=_F25_PASS)
|
||||
assert result["pass"] is False, f"{field} 共存未拦"
|
||||
assert {c["name"]: c for c in result["checks"]}["acceptedWithoutConflict"]["pass"] is False
|
||||
|
||||
|
||||
def test_shadow20_identity_drift_rejected():
|
||||
"""固定六项漂移即拒(shadow 可比性前提)。"""
|
||||
rows = _shadow_rows()
|
||||
rows[9]["identity"] = {**_IDENTITY, "commitHash": "drifted1"}
|
||||
result = evaluate_shadow20_gate(rows, fresh25_gate=_F25_PASS)
|
||||
assert result["pass"] is False
|
||||
assert {c["name"]: c for c in result["checks"]}["identityFixed"]["pass"] is False
|
||||
|
||||
|
||||
def test_shadow20_identity_drift_against_plan_rejected():
|
||||
plan = build_shadow_plan([f"prompt-{i}" for i in range(20)], commit_hash=_IDENTITY["commitHash"],
|
||||
chrome_version=_IDENTITY["chromeVersion"], actor_model=_IDENTITY["actorModel"],
|
||||
judge_model=_IDENTITY["judgeModel"], prompt_version=_IDENTITY["promptVersion"],
|
||||
config_snapshot_hash=_IDENTITY["configSnapshotHash"])
|
||||
rows = _shadow_rows()
|
||||
rows[0]["identity"] = {**_IDENTITY, "promptVersion": "3.0.13"} # 与 plan 不一致
|
||||
result = evaluate_shadow20_gate(rows, fresh25_gate=_F25_PASS, plan=plan)
|
||||
assert result["pass"] is False
|
||||
|
||||
|
||||
def test_shadow20_human_review_coverage_enforced():
|
||||
"""分歧/reject/inconclusive/tester_error 必查 + 普通 accept 至少抽 5。"""
|
||||
rows = _shadow_rows()
|
||||
rows[4]["outcome"] = "reject"
|
||||
rows[4]["humanReviewed"] = False # 必查局未复核
|
||||
result = evaluate_shadow20_gate(rows, fresh25_gate=_F25_PASS)
|
||||
assert result["pass"] is False
|
||||
assert {c["name"]: c for c in result["checks"]}["humanReviewCoverage"]["pass"] is False
|
||||
|
||||
# accept 抽查不足 5:6 个 accept 只复核 1 个
|
||||
rows2 = _shadow_rows()
|
||||
for i in range(14):
|
||||
rows2[i]["outcome"] = "reject" # 14 reject(全已复核)+ 6 accept
|
||||
for i in range(14, 19):
|
||||
rows2[i]["humanReviewed"] = False # 6 accept 中 5 个未复核 → 只 1 个 < 5
|
||||
result2 = evaluate_shadow20_gate(rows2, fresh25_gate=_F25_PASS)
|
||||
assert result2["pass"] is False
|
||||
|
||||
|
||||
def test_shadow20_plan_builder_validates_fixed_identity():
|
||||
prompts = [f"prompt-{i}" for i in range(20)]
|
||||
plan = build_shadow_plan(prompts, commit_hash="c1", chrome_version="126",
|
||||
actor_model="MiniMax-M3", judge_model="MiniMax-M3",
|
||||
prompt_version="3.0.14", config_snapshot_hash="cfg1")
|
||||
assert plan["acceptanceMode"] == "v3_shadow"
|
||||
assert len(plan["prompts"]) == 20
|
||||
with pytest.raises(ValueError, match="≥20"):
|
||||
build_shadow_plan(prompts[:19], commit_hash="c1", chrome_version="126",
|
||||
actor_model="m", judge_model="m", prompt_version="p", config_snapshot_hash="c")
|
||||
with pytest.raises(ValueError, match="不得为空"):
|
||||
build_shadow_plan(prompts, commit_hash="", chrome_version="126",
|
||||
actor_model="m", judge_model="m", prompt_version="p", config_snapshot_hash="c")
|
||||
|
||||
|
||||
def test_shadow20_runner_framework_stamps_identity_serially():
|
||||
"""runner 框架:串行调用注入的真跑 callable,每局盖固定身份戳(罐头 run_one,零模型)。"""
|
||||
plan = build_shadow_plan([f"p{i}" for i in range(20)], commit_hash="c1", chrome_version="126",
|
||||
actor_model="MiniMax-M3", judge_model="MiniMax-M3",
|
||||
prompt_version="3.0.14", config_snapshot_hash="cfg1")
|
||||
calls = []
|
||||
|
||||
async def _fake_run_one(prompt, identity):
|
||||
calls.append(prompt)
|
||||
return {"outcome": "accept", "proofComplete": True, "humanReviewed": True}
|
||||
|
||||
rows = asyncio.run(run_shadow_batch(plan, run_one=_fake_run_one))
|
||||
assert len(rows) == 20 and len(calls) == 20
|
||||
assert all(row["identity"]["commitHash"] == "c1" for row in rows)
|
||||
assert [row["gid"] for row in rows[:2]] == ["shadow20-01", "shadow20-02"]
|
||||
# 框架产出的行直接可进闸门(全绿 + fresh25 前置 → PASS)
|
||||
assert evaluate_shadow20_gate(rows, fresh25_gate=_F25_PASS, plan=plan)["pass"] is True
|
||||
|
||||
|
||||
def test_shadow20_runner_rejects_malformed_plan():
|
||||
with pytest.raises(ValueError):
|
||||
asyncio.run(run_shadow_batch({"prompts": "not-a-list"}, run_one=lambda p, i: {}))
|
||||
@ -230,6 +230,105 @@ def test_drive_cheap_generation_fake_sse(tmp_path, monkeypatch):
|
||||
assert s2["stoppedReason"] == "total_timeout"
|
||||
|
||||
|
||||
def test_drive_v3_verified_reject_posts_exactly_one_repair(tmp_path, monkeypatch):
|
||||
"""Service v3 不靠 RepairMiddleware;只有首轮 verified reject 才给同一 session 再发一次修复。"""
|
||||
import asyncio
|
||||
import httpx
|
||||
import _bootstrap
|
||||
import cheap_studio
|
||||
import cheap_verify
|
||||
import service.control_plane as CP
|
||||
|
||||
_install_common_stubs(tmp_path, monkeypatch)
|
||||
# W-GOLD-LIVE 检查点 2:编排升 acceptance-request/3,validate.py 语义层钉死 /3 配 2026-07-15.v3
|
||||
# 版本线(rolling proof-obligations.v2.json + 完整 ProofObligationRegistry schema),pin 随接线同步。
|
||||
assert cheap_verify._V3_OBLIGATIONS_FILE.name == "proof-obligations.v2.json"
|
||||
assert cheap_verify._V3_OBLIGATIONS_SCHEMA.name == "proof-obligation-registry.schema.json"
|
||||
monkeypatch.setattr(cheap_studio, "acceptance_v3_mode", lambda: "v3")
|
||||
monkeypatch.setattr(_bootstrap, "ensure_api_key_env", lambda: None)
|
||||
verdict = {"pass": True, "guards": {name: {"pass": True} for name in cheap_verify._FLOOR_GATES}}
|
||||
floor_calls = []
|
||||
|
||||
async def _floor(gid):
|
||||
floor_calls.append(gid)
|
||||
return verdict
|
||||
|
||||
monkeypatch.setattr(D, "_run_v3_floor_gates", _floor)
|
||||
svc_rows = iter([
|
||||
{"costRmb": 0.4, "rmbGate": "active", "repairs": 0},
|
||||
{"costRmb": 0.2, "rmbGate": "active", "repairs": 0},
|
||||
])
|
||||
monkeypatch.setattr(D, "_read_service_run_summary", lambda gid, wait_s=5.0: next(svc_rows))
|
||||
|
||||
acceptance_requests = []
|
||||
|
||||
async def _accept(req):
|
||||
acceptance_requests.append(req)
|
||||
if len(acceptance_requests) == 1:
|
||||
return {"runId": "r1", "floor": {"pass": True},
|
||||
"artifactHash": "a" * 64,
|
||||
"acceptanceRequestHash": req["acceptanceIdentity"]["acceptanceRequestHash"],
|
||||
"decision": {"outcome": "reject", "accepted": False, "publishFrozen": True,
|
||||
"repairEligible": True, "repairFeedback": "按硬证修一次",
|
||||
"parentChainCostRmb": 0.3},
|
||||
"compatibility": {"accepted": False, "ok": False, "acceptanceVersion": "v3",
|
||||
"playtest": {}, "judge": {}, "failureLayer": {"layer": "gameplay"},
|
||||
"failureReason": "硬证拒绝", "trace": {}}}
|
||||
return {"runId": "r2", "floor": {"pass": True},
|
||||
"artifactHash": "b" * 64,
|
||||
"acceptanceRequestHash": req["acceptanceIdentity"]["acceptanceRequestHash"],
|
||||
"decision": {"outcome": "accept", "accepted": True, "publishFrozen": False,
|
||||
"repairEligible": False, "parentChainCostRmb": 0.5},
|
||||
"compatibility": {"accepted": True, "ok": True, "acceptanceVersion": "v3",
|
||||
"playtest": {"outcome": "accept"}, "judge": {"verdict": "accept"},
|
||||
"failureLayer": {"layer": "none"}, "failureReason": None, "trace": {}}}
|
||||
|
||||
monkeypatch.setattr(cheap_verify, "run_acceptance_v3", _accept)
|
||||
monkeypatch.setattr(cheap_verify, "is_v3_repair_authorized",
|
||||
lambda payload: (payload.get("decision") or {}).get("repairEligible") is True)
|
||||
monkeypatch.setattr(CP, "_wait_for_turn_end", lambda *a, **k: None)
|
||||
|
||||
async def _ended(*a, **k):
|
||||
return {"ended": True, "reason": "REPLY_END", "endEvent": {}}
|
||||
|
||||
monkeypatch.setattr(CP, "_wait_for_turn_end", _ended)
|
||||
chat_texts = []
|
||||
|
||||
class _Resp:
|
||||
def __init__(self, data): self._data = data
|
||||
def json(self): return self._data
|
||||
|
||||
class _Http:
|
||||
def __init__(self, *a, **k): pass
|
||||
async def __aenter__(self): return self
|
||||
async def __aexit__(self, *a): return False
|
||||
async def post(self, url, **kwargs):
|
||||
if url.endswith("/chat/"):
|
||||
chat_texts.append(kwargs["json"]["input"]["content"][0]["text"])
|
||||
return _Resp({})
|
||||
if url.endswith("/credential/"): return _Resp({"credential_id": "c1"})
|
||||
if url.endswith("/agent/"): return _Resp({"agent_id": "a1"})
|
||||
if url.endswith("/sessions/"): return _Resp({"session_id": "s1"})
|
||||
return _Resp({})
|
||||
async def patch(self, *a, **k): return _Resp({})
|
||||
|
||||
monkeypatch.setattr(httpx, "AsyncClient", _Http)
|
||||
summary, _ = asyncio.run(D.drive_cheap_generation(
|
||||
{"gameId": "g-v3", "traceId": "t-v3", "brief": "解谜点击"}))
|
||||
assert len(chat_texts) == 2 and chat_texts[1].startswith("按硬证修一次")
|
||||
assert "check → build → finish" in chat_texts[1]
|
||||
assert len(floor_calls) == 2 and len(acceptance_requests) == 2
|
||||
assert acceptance_requests[1]["repairCountAcrossParentChain"] == 1
|
||||
assert acceptance_requests[1]["parentRunId"] == "r1"
|
||||
assert acceptance_requests[0]["writerCostRmb"] == 0.4
|
||||
assert acceptance_requests[1]["writerCostRmb"] == 0.2
|
||||
assert all("parentChainCostRmb" not in request for request in acceptance_requests)
|
||||
assert summary["acceptanceV3"]["runId"] == "r2" and summary["accepted"] is True
|
||||
assert summary["acceptanceV3FirstPass"]["runId"] == "r1"
|
||||
assert summary["repairAttempted"] is True and summary["acceptedAfterRepair"] is True
|
||||
assert summary["attempts"] == 2 and summary["costRmb"] == 0.6
|
||||
|
||||
|
||||
def _install_fake_http(monkeypatch, cred=None):
|
||||
"""装 fake httpx.AsyncClient:setup 三 POST 返可控体(cred 缺省给全 id),patch 空体。供回合前 unlink / fail-fast 用例复用。"""
|
||||
import httpx
|
||||
|
||||
272
cheap-worker/tests/test_full_gate.py
Normal file
272
cheap-worker/tests/test_full_gate.py
Normal file
@ -0,0 +1,272 @@
|
||||
"""full_gate 集成 runner 测试:罐头子门结果验证组合逻辑(确定性,零模型零网络)。
|
||||
|
||||
覆盖设计档 §3.2 四态降级 + §5.2 失败归因红线:全过→PASS、任一 reject→降级 reject、
|
||||
任一 tester_error/缺产物/degraded→tester_error(绝不伪装 accept,也绝不伪装 gameplay reject)、
|
||||
优先级 tester_error > reject > inconclusive > accept。
|
||||
"""
|
||||
|
||||
import sys
|
||||
from pathlib import Path
|
||||
|
||||
sys.path.insert(0, str(Path(__file__).resolve().parents[1]))
|
||||
|
||||
import cheap_verify as V # noqa: E402
|
||||
import full_gate # noqa: E402
|
||||
from full_gate import run_full_gate # noqa: E402
|
||||
|
||||
|
||||
# ── 罐头子门产物(绿态)──
|
||||
|
||||
def _green_verdict():
|
||||
return {"ok": True, "pass": True,
|
||||
"guards": {name: {"pass": True} for name in
|
||||
("A_boot", "B_uncaught", "C_frame", "D_render", "E_live",
|
||||
"F_wiring", "G_input", "H_progress", "I_control")}}
|
||||
|
||||
|
||||
def _green_floor():
|
||||
return {"accepted": True, "verdict": "accept", "rejectClasses": [], "degraded": False}
|
||||
|
||||
|
||||
def _green_v3_payload(mode="v3"):
|
||||
return {"schemaVersion": "playtest/3", "acceptanceMode": mode, "outcome": "accept",
|
||||
"decision": {"accepted": True, "publishFrozen": mode != "v3",
|
||||
"parentChainCostRmb": 1.2},
|
||||
"finalPostguard": {"pass": True, "checks": {"floorPass": True}},
|
||||
"compatibility": {"accepted": mode == "v3", "ok": mode == "v3",
|
||||
"publishFrozen": mode != "v3"}}
|
||||
|
||||
|
||||
def _green_prompt_eval():
|
||||
return {"promptId": "cheap-actor", "infrastructureComplete": True, "allGreen": True,
|
||||
"gate1_schema": True, "gate2_success": True, "gate3_regression": True,
|
||||
"gate4_cost_latency": True, "gate5_stability": True}
|
||||
|
||||
|
||||
def _noop_validator(_payload):
|
||||
return [] # 罐头 schema 校验:全绿路径不真跑 canonical validator(真校验另有契约测试覆盖)
|
||||
|
||||
|
||||
def _green_reconciler(consumptions, *, consumer_ref=None, registry=None):
|
||||
return {"ok": True, "consumed": [{"recordId": c if isinstance(c, str) else c["recordId"],
|
||||
"role": "harness_fixture", "artifactHash": "h"}
|
||||
for c in consumptions], "errors": []}
|
||||
|
||||
|
||||
def _all_green_kwargs(**overrides):
|
||||
kwargs = {
|
||||
"verdict": _green_verdict(), "floor_judgment": _green_floor(),
|
||||
"v3_payload": _green_v3_payload(), "prompt_eval_record": _green_prompt_eval(),
|
||||
"schema_validator": _noop_validator, "gold_reconciler": _green_reconciler,
|
||||
}
|
||||
kwargs.update(overrides)
|
||||
return kwargs
|
||||
|
||||
|
||||
# ── 全过 → PASS ──
|
||||
|
||||
def test_all_green_passes_and_publishable():
|
||||
result = run_full_gate(**_all_green_kwargs())
|
||||
assert result["pass"] is True
|
||||
assert result["outcome"] == "accept"
|
||||
assert result["publishable"] is True
|
||||
assert result["reasons"] == []
|
||||
assert {name: gate["state"] for name, gate in result["gates"].items()} == {
|
||||
"mechanicalNine": "pass", "visualFloor": "pass", "playtestV3": "pass",
|
||||
"promptEval": "pass", "schemaSemantics": "pass", "goldReconcile": "pass",
|
||||
}
|
||||
|
||||
|
||||
def test_shadow_accept_passes_but_not_publishable():
|
||||
"""v3_shadow 的 accept 只供校准:full_gate 可判过,但发布冻结(§3.10)。"""
|
||||
result = run_full_gate(**_all_green_kwargs(v3_payload=_green_v3_payload(mode="v3_shadow")))
|
||||
assert result["pass"] is True and result["outcome"] == "accept"
|
||||
assert result["publishable"] is False
|
||||
|
||||
|
||||
def test_vacuous_gold_reconcile_passes_without_declaration():
|
||||
"""未声明参照资产消费 ≡ 旧路径 vacuous 放行(与 _v3_check_declared_reference_assets 语义一致)。"""
|
||||
result = run_full_gate(**_all_green_kwargs())
|
||||
assert result["gates"]["goldReconcile"]["vacuous"] is True
|
||||
assert result["pass"] is True
|
||||
|
||||
|
||||
# ── 任一 reject → 降级 reject ──
|
||||
|
||||
def test_mechanical_failure_degrades_to_reject():
|
||||
verdict = _green_verdict()
|
||||
verdict["pass"] = False
|
||||
verdict["guards"]["A_boot"] = {"pass": False, "err": "boot 崩溃"}
|
||||
result = run_full_gate(**_all_green_kwargs(verdict=verdict))
|
||||
assert result["pass"] is False
|
||||
assert result["outcome"] == "reject"
|
||||
assert result["gates"]["mechanicalNine"]["failedGates"] == ["A_boot"]
|
||||
|
||||
|
||||
def test_check_level_structural_failure_is_reject():
|
||||
"""Node check 级失败(src/ 缺失,ok=False)→ 机械 reject。"""
|
||||
result = run_full_gate(**_all_green_kwargs(verdict={"ok": False, "errors": ["src/ 目录不存在或不可读"]}))
|
||||
assert result["pass"] is False and result["outcome"] == "reject"
|
||||
|
||||
|
||||
def test_visual_floor_reject_degrades_to_reject():
|
||||
floor = {"accepted": False, "verdict": "reject", "rejectClasses": ["hollow"], "degraded": False}
|
||||
result = run_full_gate(**_all_green_kwargs(floor_judgment=floor))
|
||||
assert result["pass"] is False and result["outcome"] == "reject"
|
||||
assert result["gates"]["visualFloor"]["rejectClasses"] == ["hollow"]
|
||||
|
||||
|
||||
def test_playtest_v3_reject_degrades_to_reject():
|
||||
payload = _green_v3_payload()
|
||||
payload.update({"outcome": "reject",
|
||||
"decision": {"accepted": False, "failure": {"layer": "gameplay"}}})
|
||||
payload["finalPostguard"]["pass"] = False
|
||||
result = run_full_gate(**_all_green_kwargs(v3_payload=payload))
|
||||
assert result["pass"] is False and result["outcome"] == "reject"
|
||||
|
||||
|
||||
def test_gold_reconcile_errors_are_verified_reject():
|
||||
"""声明消费而对账失败 = verified reject(设计语义)。"""
|
||||
def _bad_reconciler(consumptions, *, consumer_ref=None, registry=None):
|
||||
return {"ok": False, "consumed": [], "errors": ["参照资产未激活不得消费:ref-001(lifecycleStatus=candidate)"]}
|
||||
result = run_full_gate(**_all_green_kwargs(
|
||||
reference_consumptions=["ref-001"], gold_reconciler=_bad_reconciler))
|
||||
assert result["pass"] is False and result["outcome"] == "reject"
|
||||
assert "未激活" in result["gates"]["goldReconcile"]["reason"]
|
||||
|
||||
|
||||
# ── tester_error 绝不伪装 accept ──
|
||||
|
||||
def test_playtest_tester_error_never_disguises_accept():
|
||||
payload = _green_v3_payload()
|
||||
payload.update({"outcome": "tester_error",
|
||||
"decision": {"accepted": False, "failure": {"layer": "tester_error",
|
||||
"subtype": "environment_error"}}})
|
||||
payload["finalPostguard"]["pass"] = False
|
||||
result = run_full_gate(**_all_green_kwargs(v3_payload=payload))
|
||||
assert result["pass"] is False
|
||||
assert result["outcome"] == "tester_error" # 不是 accept,也不是 reject
|
||||
|
||||
|
||||
def test_self_contradictory_payload_is_tester_error():
|
||||
"""decision.accepted=True 但 outcome≠accept:封存产物自相矛盾 = 仪器异常,绝不放行。"""
|
||||
payload = _green_v3_payload()
|
||||
payload["outcome"] = "reject" # accepted 仍为 True → 矛盾
|
||||
result = run_full_gate(**_all_green_kwargs(v3_payload=payload))
|
||||
assert result["pass"] is False and result["outcome"] == "tester_error"
|
||||
assert "矛盾" in result["gates"]["playtestV3"]["reason"]
|
||||
|
||||
|
||||
def test_degraded_floor_is_tester_error_not_gameplay_reject():
|
||||
"""地板 degraded(评不出)按 §5.2 归仪器异常,不伪装 gameplay reject。"""
|
||||
floor = {"accepted": False, "verdict": "reject", "rejectClasses": ["degraded"],
|
||||
"degraded": True, "reason": "真玩截图证据缺失"}
|
||||
result = run_full_gate(**_all_green_kwargs(floor_judgment=floor))
|
||||
assert result["pass"] is False and result["outcome"] == "tester_error"
|
||||
|
||||
|
||||
def test_missing_artifacts_fail_closed_to_tester_error():
|
||||
"""任何子门产物缺失 → fail-closed tester_error,六个都缺也绝不 accept。"""
|
||||
result = run_full_gate(schema_validator=_noop_validator, gold_reconciler=_green_reconciler)
|
||||
assert result["pass"] is False and result["outcome"] == "tester_error"
|
||||
assert result["publishable"] is False
|
||||
missing = [name for name, gate in result["gates"].items() if gate["state"] == "tester_error"]
|
||||
assert set(missing) == {"mechanicalNine", "visualFloor", "playtestV3", "promptEval", "schemaSemantics"}
|
||||
|
||||
|
||||
def test_schema_errors_are_tester_error():
|
||||
"""canonical validate.py 报错 = schema 异常 → tester_error(§5.2)。"""
|
||||
result = run_full_gate(**_all_green_kwargs(
|
||||
schema_validator=lambda p: ["#/events/0 payloadHash 与 payloadCanonical 不一致"]))
|
||||
assert result["pass"] is False and result["outcome"] == "tester_error"
|
||||
assert result["gates"]["schemaSemantics"]["errors"]
|
||||
|
||||
|
||||
def test_prompt_eval_red_gate_is_tester_error():
|
||||
"""校准闸红 = Actor/Judge 判读不可信 → tester_error,绝不放行。"""
|
||||
record = _green_prompt_eval()
|
||||
record.update({"allGreen": False, "gate3_regression": False})
|
||||
result = run_full_gate(**_all_green_kwargs(prompt_eval_record=record))
|
||||
assert result["pass"] is False and result["outcome"] == "tester_error"
|
||||
assert "gate3_regression" in result["gates"]["promptEval"]["failedRecords"][0]["reason"]
|
||||
|
||||
|
||||
def test_prompt_eval_infrastructure_uncertain_is_tester_error():
|
||||
record = _green_prompt_eval()
|
||||
record["infrastructureComplete"] = False
|
||||
result = run_full_gate(**_all_green_kwargs(prompt_eval_record=record))
|
||||
assert result["pass"] is False and result["outcome"] == "tester_error"
|
||||
|
||||
|
||||
def test_prompt_eval_multiple_records_all_must_be_green():
|
||||
"""Actor + Judge A/B 多条记录:任一不绿整体 tester_error。"""
|
||||
judge_b = _green_prompt_eval()
|
||||
judge_b.update({"promptId": "cheap-judge-b", "allGreen": False, "gate2_success": False})
|
||||
result = run_full_gate(**_all_green_kwargs(
|
||||
prompt_eval_record=[_green_prompt_eval(), judge_b]))
|
||||
assert result["pass"] is False and result["outcome"] == "tester_error"
|
||||
assert result["gates"]["promptEval"]["failedRecords"][0]["promptId"] == "cheap-judge-b"
|
||||
|
||||
|
||||
# ── inconclusive 与优先级 ──
|
||||
|
||||
def test_playtest_inconclusive_degrades_to_inconclusive():
|
||||
payload = _green_v3_payload()
|
||||
payload.update({"outcome": "inconclusive",
|
||||
"decision": {"accepted": False, "failure": {"layer": "gameplay",
|
||||
"subtype": "proof_missing"}}})
|
||||
payload["finalPostguard"]["pass"] = False
|
||||
result = run_full_gate(**_all_green_kwargs(v3_payload=payload))
|
||||
assert result["pass"] is False and result["outcome"] == "inconclusive"
|
||||
|
||||
|
||||
def test_tester_error_outranks_reject():
|
||||
"""一子门 reject + 另一子门 tester_error → 整体 tester_error(仪器异常时不能宣称已证缺陷)。"""
|
||||
floor = {"accepted": False, "verdict": "reject", "rejectClasses": ["broken"], "degraded": False}
|
||||
payload = _green_v3_payload()
|
||||
payload.update({"outcome": "reject", "decision": {"accepted": False}})
|
||||
payload["finalPostguard"]["pass"] = False
|
||||
result = run_full_gate(**_all_green_kwargs(floor_judgment=floor, v3_payload=payload,
|
||||
schema_validator=lambda p: ["schema 脏"]))
|
||||
assert result["pass"] is False
|
||||
assert result["outcome"] == "tester_error" # schema 异常压过 floor reject
|
||||
|
||||
|
||||
def test_reject_outranks_inconclusive():
|
||||
payload = _green_v3_payload()
|
||||
payload.update({"outcome": "inconclusive", "decision": {"accepted": False}})
|
||||
payload["finalPostguard"]["pass"] = False
|
||||
floor = {"accepted": False, "verdict": "reject", "rejectClasses": ["off_brief"], "degraded": False}
|
||||
result = run_full_gate(**_all_green_kwargs(v3_payload=payload, floor_judgment=floor))
|
||||
assert result["pass"] is False and result["outcome"] == "reject"
|
||||
|
||||
|
||||
# ── 默认绑定真子门函数(消费不重写)──
|
||||
|
||||
def test_default_validators_bind_real_subgate_functions():
|
||||
"""不注入时默认绑 cheap_verify 真函数:canonical 校验 + 金标对账,保证集成 runner 消费子门而非影子实现。"""
|
||||
import inspect
|
||||
src = inspect.getsource(full_gate.run_full_gate)
|
||||
assert "V.validate_acceptance_v3_payload" in src
|
||||
assert "V.reconcile_v3_reference_asset_consumption" in src
|
||||
# 默认 schema_validator 对非法 payload 必须 fail-closed 报错(真跑 canonical 校验器)。
|
||||
result = run_full_gate(verdict=_green_verdict(), floor_judgment=_green_floor(),
|
||||
v3_payload={"schemaVersion": "playtest/3"},
|
||||
prompt_eval_record=_green_prompt_eval())
|
||||
assert result["gates"]["schemaSemantics"]["state"] == "tester_error"
|
||||
assert result["gates"]["schemaSemantics"]["errors"]
|
||||
|
||||
|
||||
def test_gold_reconcile_default_uses_real_registry_when_declared():
|
||||
"""声明消费 + 默认注册表(当前全是 migration_pending、无 active)→ 对账拒绝 reject。
|
||||
|
||||
这正是 W-GOLD-LIVE「完成前不得新增 live 消费」的机器强制点:声明消费一条在册但未激活的记录即被拒。
|
||||
"""
|
||||
result = run_full_gate(verdict=_green_verdict(), floor_judgment=_green_floor(),
|
||||
v3_payload=_green_v3_payload(), prompt_eval_record=_green_prompt_eval(),
|
||||
schema_validator=_noop_validator,
|
||||
reference_consumptions=["gold-m3-gem-r3"])
|
||||
assert result["gates"]["goldReconcile"]["state"] == "reject"
|
||||
assert result["gates"]["goldReconcile"]["vacuous"] is False
|
||||
assert "未激活" in result["gates"]["goldReconcile"]["reason"]
|
||||
assert result["outcome"] == "reject"
|
||||
1104
cheap-worker/tests/test_reference_asset_gate.py
Normal file
1104
cheap-worker/tests/test_reference_asset_gate.py
Normal file
File diff suppressed because it is too large
Load Diff
405
cheap-worker/tests/test_reference_asset_service_wiring.py
Normal file
405
cheap-worker/tests/test_reference_asset_service_wiring.py
Normal file
@ -0,0 +1,405 @@
|
||||
"""参照资产 v2 在 CLI/Service 生产入口的冻结快照接线回归。"""
|
||||
|
||||
import asyncio
|
||||
import errno
|
||||
import json
|
||||
import shutil
|
||||
import sys
|
||||
from pathlib import Path
|
||||
|
||||
import pytest
|
||||
|
||||
# brief 的验证命令从仓根启动;显式加入 cheap-worker,保持与其它 Service 测试一致。
|
||||
sys.path.insert(0, str(Path(__file__).resolve().parents[1]))
|
||||
|
||||
import cheap_run
|
||||
import cheap_service_app as A
|
||||
import cheap_service_driver as D
|
||||
import cheap_verify
|
||||
import reference_asset_gate
|
||||
|
||||
|
||||
def _fake_verified():
|
||||
"""构造最小只读验证结果;hash 均由真实 canonical 实现计算。"""
|
||||
files = {
|
||||
"assets/gold/guide.md": b"captured guide",
|
||||
"assets/gold/sub/rules.txt": b"captured rules",
|
||||
}
|
||||
snapshot_hash = reference_asset_gate.snapshot_hash(files)
|
||||
receipt = {
|
||||
"schemaVersion": "ReferenceAssetVerificationReceipt/1",
|
||||
"receiptId": "receipt-survivor-gold-v1-gac-shanhai-xingji-test",
|
||||
"recordId": "gac-shanhai-xingji",
|
||||
"finalSnapshotHash": snapshot_hash,
|
||||
}
|
||||
return reference_asset_gate.VerifiedReferenceAssets(
|
||||
constraint_records=({"recordId": "gac-shanhai-xingji", "role": "game_content_gold"},),
|
||||
files=files,
|
||||
receipts=(receipt,),
|
||||
snapshot_hash=snapshot_hash,
|
||||
reference_roots={"gac-shanhai-xingji": ("assets/gold",)},
|
||||
)
|
||||
|
||||
|
||||
def _cfg_path(tmp_path: Path, session_id: str) -> Path:
|
||||
return tmp_path / "_cheap-sessions" / f"{session_id}.json"
|
||||
|
||||
|
||||
def _snapshot_dir(cfg_path: Path) -> Path:
|
||||
return cfg_path.with_name(f"{cfg_path.stem}.reference-assets")
|
||||
|
||||
|
||||
def _tool_function(tools, name: str):
|
||||
"""从 Service 工厂返回的真实 FunctionTool 列表取目标函数。"""
|
||||
for tool in tools:
|
||||
if tool.name == name:
|
||||
return tool._func
|
||||
raise AssertionError(f"Service Toolkit 缺少工具:{name}")
|
||||
|
||||
|
||||
def _publish(tmp_path, monkeypatch, session_id="session-reference"):
|
||||
"""通过 driver 真实物化 sidecar 与独立 session snapshot。"""
|
||||
monkeypatch.setattr(cheap_run, "session_cfg_path", lambda sid: _cfg_path(tmp_path, sid))
|
||||
verified = _fake_verified()
|
||||
assert D._write_session_cfg(
|
||||
session_id,
|
||||
external_game_id="70012",
|
||||
reference_assets=verified,
|
||||
) is True
|
||||
return _cfg_path(tmp_path, session_id), verified
|
||||
|
||||
|
||||
def test_service_factory_reads_only_driver_materialized_snapshot(tmp_path, monkeypatch):
|
||||
"""工厂只读 session snapshot;活目录后续改写不得影响 read/list。"""
|
||||
cfg_path, _ = _publish(tmp_path, monkeypatch)
|
||||
live_calls = []
|
||||
|
||||
def live_read(path):
|
||||
live_calls.append(("read", path))
|
||||
return {"ok": True, "content": "rewritten live asset", "truncated": False}
|
||||
|
||||
def live_list(path):
|
||||
live_calls.append(("list", path))
|
||||
return {"ok": True, "entries": ["rewritten.txt"]}
|
||||
|
||||
monkeypatch.setattr(cheap_run, "read_file", live_read)
|
||||
monkeypatch.setattr(cheap_run, "list_dir", live_list)
|
||||
tools = asyncio.run(A._cheap_tools_factory("u", "a", "session-reference"))
|
||||
|
||||
read = asyncio.run(_tool_function(tools, "read_file")(path="assets/gold/guide.md"))
|
||||
listed = asyncio.run(_tool_function(tools, "list_dir")(path="assets/gold"))
|
||||
assert read == "captured guide"
|
||||
assert listed == "guide.md\nsub/"
|
||||
assert live_calls == []
|
||||
|
||||
# sidecar 只允许稳定策略元数据与文件索引,不落 record/identity 自报 artifact hash。
|
||||
cfg_text = cfg_path.read_text(encoding="utf-8")
|
||||
cfg = json.loads(cfg_text)
|
||||
policy = cfg["reference_asset_policy"]
|
||||
assert policy["policy_id"] == "survivor-gold-v1"
|
||||
assert policy["mode"] == "frozen_preflight"
|
||||
assert "artifactHash" not in cfg_text
|
||||
assert {item["path"] for item in policy["files"]} == {
|
||||
"assets/gold/guide.md", "assets/gold/sub/rules.txt"}
|
||||
|
||||
|
||||
@pytest.mark.parametrize("mutation", ["missing_policy", "bad_json", "non_object"])
|
||||
def test_service_factory_rejects_existing_snapshot_without_valid_policy(
|
||||
tmp_path, monkeypatch, mutation):
|
||||
"""已有冻结 snapshot 缺 policy 或配置损坏时必须失败,绝不回读活目录。"""
|
||||
cfg_path, _ = _publish(tmp_path, monkeypatch, session_id=f"cfg-{mutation}")
|
||||
if mutation == "missing_policy":
|
||||
cfg = json.loads(cfg_path.read_text(encoding="utf-8"))
|
||||
del cfg["reference_asset_policy"]
|
||||
cfg_path.write_text(json.dumps(cfg, ensure_ascii=False), encoding="utf-8")
|
||||
elif mutation == "bad_json":
|
||||
cfg_path.write_text("{broken", encoding="utf-8")
|
||||
else:
|
||||
cfg_path.write_text("[]", encoding="utf-8")
|
||||
|
||||
live_calls = []
|
||||
|
||||
def live_read(path):
|
||||
live_calls.append(("read", path))
|
||||
raise AssertionError("冻结 snapshot 配置异常时不得读取活目录")
|
||||
|
||||
monkeypatch.setattr(cheap_run, "read_file", live_read)
|
||||
with pytest.raises(ValueError):
|
||||
asyncio.run(A._cheap_tools_factory("u", "a", f"cfg-{mutation}"))
|
||||
assert live_calls == []
|
||||
|
||||
|
||||
def test_existing_snapshot_rejects_default_sidecar_overwrite_without_live_read(
|
||||
tmp_path, monkeypatch):
|
||||
"""同 session 的默认写入不得覆盖冻结 policy,工厂仍只能读原 snapshot。"""
|
||||
cfg_path, _ = _publish(tmp_path, monkeypatch, session_id="reuse")
|
||||
|
||||
assert D._write_session_cfg("reuse", external_game_id="70099") is False
|
||||
cfg = json.loads(cfg_path.read_text(encoding="utf-8"))
|
||||
assert cfg["external_game_id"] == "70012"
|
||||
assert "reference_asset_policy" in cfg
|
||||
|
||||
live_calls = []
|
||||
monkeypatch.setattr(
|
||||
cheap_run, "read_file",
|
||||
lambda path: live_calls.append(("read", path)) or "live")
|
||||
tools = asyncio.run(A._cheap_tools_factory("u", "a", "reuse"))
|
||||
assert asyncio.run(_tool_function(tools, "read_file")(path="assets/gold/guide.md")) == "captured guide"
|
||||
assert live_calls == []
|
||||
|
||||
|
||||
def test_missing_snapshot_rejects_default_sidecar_overwrite_of_frozen_cfg(
|
||||
tmp_path, monkeypatch):
|
||||
"""snapshot 被删除后,默认写入仍不得覆盖原冻结 cfg 或回落活目录。"""
|
||||
cfg_path, _ = _publish(tmp_path, monkeypatch, session_id="orphan-cfg")
|
||||
original_bytes = cfg_path.read_bytes()
|
||||
original_policy = json.loads(original_bytes)["reference_asset_policy"]
|
||||
shutil.rmtree(_snapshot_dir(cfg_path))
|
||||
|
||||
assert D._write_session_cfg(
|
||||
"orphan-cfg", external_game_id="70099", reference_assets=None) is False
|
||||
assert cfg_path.read_bytes() == original_bytes
|
||||
assert json.loads(cfg_path.read_bytes())["reference_asset_policy"] == original_policy
|
||||
|
||||
|
||||
def test_session_cfg_reader_rejects_orphan_snapshot_without_cfg(tmp_path, monkeypatch):
|
||||
"""cfg 被删除但冻结 snapshot 孤立存在时,reader 必须拒绝活目录回落。"""
|
||||
cfg_path, _ = _publish(tmp_path, monkeypatch, session_id="orphan-snapshot")
|
||||
cfg_path.unlink()
|
||||
|
||||
with pytest.raises(ValueError, match="session-cfg 缺失"):
|
||||
A._read_session_cfg("orphan-snapshot")
|
||||
|
||||
|
||||
@pytest.mark.parametrize(
|
||||
"mutation",
|
||||
["missing_dir", "missing_file", "special_file", "index", "receipt", "snapshot_hash"],
|
||||
)
|
||||
def test_service_factory_rejects_snapshot_sidecar_drift_before_tools(
|
||||
tmp_path, monkeypatch, mutation):
|
||||
"""目录、文件、特殊文件、索引、receipt 或 snapshot 任一漂移都必须 fail-closed。"""
|
||||
cfg_path, _ = _publish(tmp_path, monkeypatch, session_id=f"tamper-{mutation}")
|
||||
snapshot_dir = _snapshot_dir(cfg_path)
|
||||
target = snapshot_dir / "assets/gold/guide.md"
|
||||
|
||||
if mutation == "missing_dir":
|
||||
shutil.rmtree(snapshot_dir)
|
||||
elif mutation == "missing_file":
|
||||
target.unlink()
|
||||
elif mutation == "special_file":
|
||||
target.unlink()
|
||||
target.symlink_to(snapshot_dir / "assets/gold/sub/rules.txt")
|
||||
elif mutation == "receipt":
|
||||
receipt = snapshot_dir / ".reference-receipts.json"
|
||||
receipt.write_bytes(receipt.read_bytes() + b" ")
|
||||
else:
|
||||
cfg = json.loads(cfg_path.read_text(encoding="utf-8"))
|
||||
policy = cfg["reference_asset_policy"]
|
||||
if mutation == "index":
|
||||
policy["files"][0]["size"] += 1
|
||||
else:
|
||||
policy["snapshot_hash"] = "0" * 64
|
||||
cfg_path.write_text(json.dumps(cfg, ensure_ascii=False), encoding="utf-8")
|
||||
|
||||
with pytest.raises(ValueError):
|
||||
asyncio.run(A._cheap_tools_factory("u", "a", f"tamper-{mutation}"))
|
||||
|
||||
|
||||
def test_default_sidecar_does_not_inject_reference_snapshot(tmp_path, monkeypatch):
|
||||
"""未声明 policy 的既有 sidecar 保持活目录 read/list 默认语义。"""
|
||||
monkeypatch.setattr(cheap_run, "session_cfg_path", lambda sid: _cfg_path(tmp_path, sid))
|
||||
assert D._write_session_cfg("default", external_game_id="70013") is True
|
||||
monkeypatch.setattr(cheap_run, "read_file",
|
||||
lambda path: {"ok": True, "content": "live", "truncated": False})
|
||||
tools = asyncio.run(A._cheap_tools_factory("u", "a", "default"))
|
||||
|
||||
assert asyncio.run(_tool_function(tools, "read_file")(path="assets/gold/guide.md")) == "live"
|
||||
assert not _snapshot_dir(_cfg_path(tmp_path, "default")).exists()
|
||||
|
||||
|
||||
def test_session_cfg_reader_rejects_symlink_and_oversize(tmp_path, monkeypatch):
|
||||
"""会话配置只接受固定上限内的普通文件,拒绝 symlink 与超限输入。"""
|
||||
monkeypatch.setattr(cheap_run, "session_cfg_path", lambda sid: _cfg_path(tmp_path, sid))
|
||||
target = tmp_path / "target.json"
|
||||
target.write_text("{}", encoding="utf-8")
|
||||
symlink = _cfg_path(tmp_path, "linked")
|
||||
symlink.parent.mkdir(parents=True)
|
||||
symlink.symlink_to(target)
|
||||
with pytest.raises(ValueError):
|
||||
A._read_session_cfg("linked")
|
||||
|
||||
oversized = _cfg_path(tmp_path, "oversized")
|
||||
oversized.write_bytes(b" " * (A._SESSION_CFG_MAX_BYTES + 1))
|
||||
with pytest.raises(ValueError):
|
||||
A._read_session_cfg("oversized")
|
||||
|
||||
|
||||
def test_session_cfg_reader_rejects_symlink_swap_after_lstat(tmp_path, monkeypatch):
|
||||
"""lstat 后被换成 symlink 时,O_NOFOLLOW 拒绝不得降级成缺 sidecar。"""
|
||||
monkeypatch.setattr(cheap_run, "session_cfg_path", lambda sid: _cfg_path(tmp_path, sid))
|
||||
cfg = _cfg_path(tmp_path, "raced")
|
||||
cfg.parent.mkdir(parents=True)
|
||||
cfg.write_text("{}", encoding="utf-8")
|
||||
|
||||
def raced_open(*args, **kwargs):
|
||||
raise OSError(errno.ELOOP, "symlink swap")
|
||||
|
||||
monkeypatch.setattr(A.os, "open", raced_open)
|
||||
with pytest.raises(ValueError, match="拒绝"):
|
||||
A._read_session_cfg("raced")
|
||||
|
||||
|
||||
class _StopAtChat(BaseException):
|
||||
"""测试仅用于在首个 /chat 观察点停止真实 driver。"""
|
||||
|
||||
|
||||
def _install_driver_setup(monkeypatch, tmp_path, events, *, sidecar_ok=True,
|
||||
real_sidecar=False):
|
||||
"""隔离 Service 外部 I/O,仅保留 driver 内部发布顺序。"""
|
||||
import _bootstrap
|
||||
import cheap_otlp_sink
|
||||
import httpx
|
||||
|
||||
monkeypatch.setenv("TIER2_GEN__ACCEPTANCE__MODE", "v3_shadow")
|
||||
monkeypatch.setattr(cheap_run, "archive_prior_run", lambda game_id: None)
|
||||
monkeypatch.setattr(cheap_run, "scaffold", lambda *args, **kwargs: {"ok": True, "output": ""})
|
||||
monkeypatch.setattr(cheap_run, "clean_stale_evidence", lambda game_id: None)
|
||||
monkeypatch.setattr(cheap_run, "game_dir", lambda game_id: tmp_path / f"amgen-{game_id}")
|
||||
monkeypatch.setattr(_bootstrap, "ensure_api_key_env", lambda: None)
|
||||
monkeypatch.setattr(cheap_otlp_sink, "current_traceparent_carrier", lambda: {})
|
||||
monkeypatch.setattr(D, "_cheap_credential_payload", lambda token=None: {"data": {}})
|
||||
monkeypatch.setattr(
|
||||
cheap_verify,
|
||||
"preflight_reference_asset_policy",
|
||||
lambda policy_id, mode: _fake_verified(),
|
||||
)
|
||||
|
||||
def publish(*args, **kwargs):
|
||||
events.append("sidecar")
|
||||
return sidecar_ok
|
||||
|
||||
if real_sidecar:
|
||||
monkeypatch.setattr(cheap_run, "session_cfg_path", lambda sid: _cfg_path(tmp_path, sid))
|
||||
else:
|
||||
monkeypatch.setattr(D, "_write_session_cfg", publish)
|
||||
|
||||
class Response:
|
||||
def __init__(self, value):
|
||||
self.value = value
|
||||
|
||||
def json(self):
|
||||
return self.value
|
||||
|
||||
class Http:
|
||||
def __init__(self, *args, **kwargs):
|
||||
pass
|
||||
|
||||
async def __aenter__(self):
|
||||
return self
|
||||
|
||||
async def __aexit__(self, *args):
|
||||
return False
|
||||
|
||||
async def post(self, url, **kwargs):
|
||||
if url.endswith("/credential/"):
|
||||
return Response({"credential_id": "c1"})
|
||||
if url.endswith("/agent/"):
|
||||
return Response({"agent_id": "a1"})
|
||||
if url.endswith("/sessions/"):
|
||||
return Response({"session_id": "s1"})
|
||||
if url.endswith("/chat/"):
|
||||
events.append("chat")
|
||||
raise _StopAtChat()
|
||||
return Response({})
|
||||
|
||||
async def patch(self, *args, **kwargs):
|
||||
events.append("patch")
|
||||
return Response({})
|
||||
|
||||
monkeypatch.setattr(httpx, "AsyncClient", Http)
|
||||
|
||||
|
||||
def test_driver_publishes_reference_sidecar_before_chat(tmp_path, monkeypatch):
|
||||
"""显式 policy 必须在首个 Writer /chat 前发布 sidecar。"""
|
||||
events = []
|
||||
_install_driver_setup(monkeypatch, tmp_path, events)
|
||||
|
||||
with pytest.raises(_StopAtChat):
|
||||
asyncio.run(D.drive_cheap_generation({
|
||||
"gameId": "70014",
|
||||
"traceId": "trace-reference",
|
||||
"brief": "解谜点击",
|
||||
"referenceAssetPolicyId": "survivor-gold-v1",
|
||||
}))
|
||||
|
||||
assert events.index("sidecar") < events.index("chat")
|
||||
|
||||
|
||||
def test_driver_does_not_chat_when_reference_sidecar_publish_fails(tmp_path, monkeypatch):
|
||||
"""显式 policy 的 sidecar 原子发布失败后不得向 Writer 发送 /chat。"""
|
||||
events = []
|
||||
_install_driver_setup(monkeypatch, tmp_path, events, sidecar_ok=False)
|
||||
|
||||
summary, _ = asyncio.run(D.drive_cheap_generation({
|
||||
"gameId": "70015",
|
||||
"traceId": "trace-reference-fail",
|
||||
"brief": "解谜点击",
|
||||
"referenceAssetPolicyId": "survivor-gold-v1",
|
||||
}))
|
||||
|
||||
assert summary["ok"] is False
|
||||
assert "session snapshot" in summary["stoppedReason"]
|
||||
assert "chat" not in events
|
||||
|
||||
|
||||
def test_driver_stops_before_chat_when_frozen_cfg_is_unconfirmable(
|
||||
tmp_path, monkeypatch):
|
||||
"""默认 sidecar 无法确认既有冻结 cfg 时,driver 必须停在首个 /chat 前。"""
|
||||
events = []
|
||||
_install_driver_setup(monkeypatch, tmp_path, events, real_sidecar=True)
|
||||
cfg_path, _ = _publish(tmp_path, monkeypatch, session_id="s1")
|
||||
original_bytes = cfg_path.read_bytes()
|
||||
shutil.rmtree(_snapshot_dir(cfg_path))
|
||||
|
||||
summary, _ = asyncio.run(D.drive_cheap_generation({
|
||||
"gameId": "70017",
|
||||
"traceId": "trace-reference-orphan-cfg",
|
||||
"brief": "解谜点击",
|
||||
}))
|
||||
|
||||
assert summary["ok"] is False
|
||||
assert "reference asset" in summary["stoppedReason"]
|
||||
assert "chat" not in events
|
||||
assert cfg_path.read_bytes() == original_bytes
|
||||
|
||||
|
||||
def test_driver_cleans_real_snapshot_when_cfg_replace_fails(tmp_path, monkeypatch):
|
||||
"""真实 sidecar 第二次 os.replace 失败时必须清理 snapshot/临时文件且不 /chat。"""
|
||||
events = []
|
||||
_install_driver_setup(monkeypatch, tmp_path, events, real_sidecar=True)
|
||||
real_replace = D.os.replace
|
||||
replace_calls = []
|
||||
|
||||
def fail_cfg_replace(source, destination):
|
||||
replace_calls.append((source, destination))
|
||||
if len(replace_calls) == 2:
|
||||
raise OSError("injected cfg publish failure")
|
||||
return real_replace(source, destination)
|
||||
|
||||
monkeypatch.setattr(D.os, "replace", fail_cfg_replace)
|
||||
summary, _ = asyncio.run(D.drive_cheap_generation({
|
||||
"gameId": "70016",
|
||||
"traceId": "trace-reference-real-fail",
|
||||
"brief": "解谜点击",
|
||||
"referenceAssetPolicyId": "survivor-gold-v1",
|
||||
}))
|
||||
|
||||
cfg_path = _cfg_path(tmp_path, "s1")
|
||||
assert len(replace_calls) == 2
|
||||
assert summary["ok"] is False
|
||||
assert "session snapshot" in summary["stoppedReason"]
|
||||
assert "chat" not in events
|
||||
assert not _snapshot_dir(cfg_path).exists()
|
||||
assert not cfg_path.exists()
|
||||
assert list(cfg_path.parent.glob(f".{cfg_path.name}.*")) == []
|
||||
assert list(cfg_path.parent.glob(f".{cfg_path.stem}.reference-assets.*")) == []
|
||||
@ -7,12 +7,17 @@ test_toolkit.py — U2 工具底座(read/list/write)+ 形状门三态。
|
||||
跑:cheap-worker/.venv/bin/python cheap-worker/tests/test_toolkit.py
|
||||
"""
|
||||
|
||||
import asyncio
|
||||
import subprocess
|
||||
import sys
|
||||
from pathlib import Path
|
||||
|
||||
import pytest
|
||||
|
||||
sys.path.insert(0, str(Path(__file__).resolve().parents[1])) # → cheap-worker/
|
||||
import cheap_run
|
||||
import cheap_toolkit
|
||||
from cheap_toolkit import CheapSession, build_toolkit
|
||||
|
||||
_GAME_RUNTIME = cheap_run._GAME_RUNTIME
|
||||
_AMODEL_GEN = cheap_run._AMODEL_GEN
|
||||
@ -114,6 +119,307 @@ def test_check_and_build_clean_pass():
|
||||
assert rb["ok"], "干净源应 build PASS,实际 FAIL:\n" + rb["output"]
|
||||
|
||||
|
||||
def _tool_function(session: CheapSession, name: str):
|
||||
"""从真实 Toolkit 取出指定函数,测试仍通过 AgentScope 的生产组装路径。"""
|
||||
toolkit = build_toolkit(session)
|
||||
for group in toolkit.tool_groups:
|
||||
for tool in group.tools:
|
||||
if tool.name == name:
|
||||
return tool._func
|
||||
raise AssertionError(f"Toolkit 缺少工具:{name}")
|
||||
|
||||
|
||||
def _call_tool(session: CheapSession, name: str, **kwargs):
|
||||
"""同步测试辅助:执行生产 Toolkit 中的 async 工具函数。"""
|
||||
return asyncio.run(_tool_function(session, name)(**kwargs))
|
||||
|
||||
|
||||
def test_reference_snapshot_read_ignores_rewritten_disk(monkeypatch, tmp_path):
|
||||
"""快照捕获后磁盘同路径改写,read_file 仍只能返回旧 bytes。"""
|
||||
disk_file = tmp_path / "assets" / "gold" / "guide.md"
|
||||
disk_file.parent.mkdir(parents=True)
|
||||
disk_file.write_text("captured", encoding="utf-8")
|
||||
session = CheapSession(
|
||||
game_id="snapshot-read",
|
||||
reference_files={"assets/gold/guide.md": b"captured"},
|
||||
reference_roots={"gold": ("assets/gold",)},
|
||||
)
|
||||
disk_file.write_text("rewritten", encoding="utf-8")
|
||||
calls = []
|
||||
|
||||
def live_read(path):
|
||||
calls.append(path)
|
||||
return {"ok": True, "content": disk_file.read_text(encoding="utf-8"), "truncated": False}
|
||||
|
||||
monkeypatch.setattr(cheap_toolkit.cheap_run, "read_file", live_read)
|
||||
|
||||
assert _call_tool(session, "read_file", path="assets/gold/guide.md") == "captured"
|
||||
assert calls == []
|
||||
|
||||
|
||||
def test_reference_snapshot_rejects_unsigned_read_without_io(monkeypatch):
|
||||
"""受保护根内未列入快照的文件必须拒绝,且不得触发活目录读取。"""
|
||||
session = CheapSession(
|
||||
game_id="snapshot-read-deny",
|
||||
reference_files={"assets/gold/allowed.txt": b"allowed"},
|
||||
reference_roots={"gold": ("assets/gold",)},
|
||||
)
|
||||
calls = []
|
||||
|
||||
def live_read(path):
|
||||
calls.append(path)
|
||||
return {"ok": True, "content": "unsigned", "truncated": False}
|
||||
|
||||
monkeypatch.setattr(cheap_toolkit.cheap_run, "read_file", live_read)
|
||||
|
||||
result = _call_tool(session, "read_file", path="assets/gold/unsigned.txt")
|
||||
assert result.startswith("ERROR:")
|
||||
assert "unsigned" not in result
|
||||
assert calls == []
|
||||
|
||||
|
||||
def test_reference_snapshot_lists_only_virtual_children_without_io(monkeypatch):
|
||||
"""受保护目录只投影快照中的直接子项和虚拟子目录,不读取磁盘新增项。"""
|
||||
session = CheapSession(
|
||||
game_id="snapshot-list",
|
||||
reference_files={
|
||||
"assets/gold/allowed.txt": b"allowed",
|
||||
"assets/gold/sub/child.txt": b"child",
|
||||
},
|
||||
reference_roots={"gold": ("assets/gold",)},
|
||||
)
|
||||
calls = []
|
||||
|
||||
def live_list(path):
|
||||
calls.append(path)
|
||||
return {"ok": True, "entries": ["unsigned.txt"]}
|
||||
|
||||
monkeypatch.setattr(cheap_toolkit.cheap_run, "list_dir", live_list)
|
||||
|
||||
assert _call_tool(session, "list_dir", path="assets/gold") == "allowed.txt\nsub/"
|
||||
assert _call_tool(session, "list_dir", path="assets/gold/sub") == "child.txt"
|
||||
assert calls == []
|
||||
|
||||
|
||||
def test_reference_prefix_similar_path_keeps_live_read_semantics(monkeypatch):
|
||||
"""路径段相似但不属于受保护根时,仍走既有 cheap_run 读取语义。"""
|
||||
session = CheapSession(
|
||||
game_id="snapshot-prefix",
|
||||
reference_files={"assets/gold/allowed.txt": b"snapshot"},
|
||||
reference_roots={"gold": ("assets/gold",)},
|
||||
)
|
||||
calls = []
|
||||
|
||||
def live_read(path):
|
||||
calls.append(path)
|
||||
return {"ok": True, "content": "live-golden", "truncated": False}
|
||||
|
||||
monkeypatch.setattr(cheap_toolkit.cheap_run, "read_file", live_read)
|
||||
|
||||
assert _call_tool(session, "read_file", path="assets/golden/allowed.txt") == "live-golden"
|
||||
assert calls == ["assets/golden/allowed.txt"]
|
||||
|
||||
|
||||
def test_reference_path_escape_cannot_bypass_protected_root(monkeypatch):
|
||||
"""回到仓根的路径别名也必须拒绝,不能绕过受保护根读取活目录。"""
|
||||
session = CheapSession(
|
||||
game_id="snapshot-escape",
|
||||
reference_files={"assets/gold/allowed.txt": b"snapshot"},
|
||||
reference_roots={"gold": ("assets/gold",)},
|
||||
)
|
||||
calls = []
|
||||
|
||||
def live_read(path):
|
||||
calls.append(path)
|
||||
return {"ok": True, "content": "disk", "truncated": False}
|
||||
|
||||
monkeypatch.setattr(cheap_toolkit.cheap_run, "read_file", live_read)
|
||||
|
||||
result = _call_tool(session, "read_file", path="../games-development-ai/assets/gold/unsigned.txt")
|
||||
assert result == "ERROR: 路径必须是仓内相对路径"
|
||||
assert calls == []
|
||||
|
||||
|
||||
def test_empty_reference_snapshot_preserves_existing_read_and_list(monkeypatch):
|
||||
"""未提供快照时保持现有 read_file/list_dir 回落行为。"""
|
||||
session = CheapSession(game_id="snapshot-empty")
|
||||
read_calls = []
|
||||
list_calls = []
|
||||
|
||||
def live_read(path):
|
||||
read_calls.append(path)
|
||||
return {"ok": True, "content": "live-content", "truncated": False}
|
||||
|
||||
def live_list(path):
|
||||
list_calls.append(path)
|
||||
return {"ok": True, "entries": ["live.txt"]}
|
||||
|
||||
monkeypatch.setattr(cheap_toolkit.cheap_run, "read_file", live_read)
|
||||
monkeypatch.setattr(cheap_toolkit.cheap_run, "list_dir", live_list)
|
||||
|
||||
assert _call_tool(session, "read_file", path=".agents/skills/example.md") == "live-content"
|
||||
assert _call_tool(session, "list_dir", path=".agents/skills") == "live.txt"
|
||||
assert read_calls == [".agents/skills/example.md"]
|
||||
assert list_calls == [".agents/skills"]
|
||||
|
||||
|
||||
def test_reference_snapshot_read_keeps_200kb_truncation_marker(monkeypatch):
|
||||
"""快照 bytes 仍按现有 200KB 规则解码并带截断标记。"""
|
||||
session = CheapSession(
|
||||
game_id="snapshot-truncate",
|
||||
reference_files={"assets/gold/large.txt": b"x" * (200 * 1024 + 1)},
|
||||
reference_roots={"gold": ("assets/gold",)},
|
||||
)
|
||||
monkeypatch.setattr(
|
||||
cheap_toolkit.cheap_run,
|
||||
"read_file",
|
||||
lambda path: (_ for _ in ()).throw(AssertionError("受保护快照不应调用 cheap_run.read_file")),
|
||||
)
|
||||
|
||||
result = _call_tool(session, "read_file", path="assets/gold/large.txt")
|
||||
assert result.startswith("[内容已截断]\n")
|
||||
assert len(result) == 30000
|
||||
|
||||
|
||||
def test_reference_session_snapshot_fields_are_read_only():
|
||||
"""Session 保存的快照映射和受保护根索引不能被调用方改写。"""
|
||||
session = CheapSession(
|
||||
game_id="snapshot-immutable",
|
||||
reference_files={"assets/gold/a.txt": b"a"},
|
||||
reference_roots={"gold": ("assets/gold",)},
|
||||
)
|
||||
|
||||
with pytest.raises(TypeError):
|
||||
session.reference_files["assets/gold/b.txt"] = b"b"
|
||||
with pytest.raises(TypeError):
|
||||
session.reference_roots["other"] = ("assets/other",)
|
||||
with pytest.raises(AttributeError):
|
||||
session.reference_files = {}
|
||||
assert session.reference_roots["gold"] == ("assets/gold",)
|
||||
|
||||
|
||||
def test_reference_snapshot_requires_roots_for_nonempty_files():
|
||||
"""非空快照没有有效受保护根时,构造必须 fail-closed。"""
|
||||
for roots in (None, {}, {"gold": ()}):
|
||||
with pytest.raises((TypeError, ValueError)) as exc_info:
|
||||
CheapSession(
|
||||
game_id="snapshot-missing-roots",
|
||||
reference_files={"assets/gold/guide.md": b"secret"},
|
||||
reference_roots=roots,
|
||||
)
|
||||
assert "secret" not in str(exc_info.value)
|
||||
assert "/" not in str(exc_info.value)
|
||||
|
||||
|
||||
@pytest.mark.parametrize(
|
||||
"root",
|
||||
("", "../sensitive-root", "/absolute/sensitive-root"),
|
||||
)
|
||||
def test_reference_roots_reject_invalid_paths_without_leaking_input(root):
|
||||
"""受保护根含空、越界或绝对路径时,构造不得静默过滤或泄露输入。"""
|
||||
with pytest.raises((TypeError, ValueError)) as exc_info:
|
||||
CheapSession(
|
||||
game_id="snapshot-invalid-root",
|
||||
reference_roots={"gold": (root,)},
|
||||
)
|
||||
|
||||
message = str(exc_info.value)
|
||||
assert "sensitive-root" not in message
|
||||
assert "/absolute" not in message
|
||||
|
||||
|
||||
def test_reference_snapshot_rejects_files_outside_all_roots():
|
||||
"""任一快照文件未被受保护根覆盖时,构造必须拒绝整个索引。"""
|
||||
with pytest.raises((TypeError, ValueError)) as exc_info:
|
||||
CheapSession(
|
||||
game_id="snapshot-uncovered-file",
|
||||
reference_files={
|
||||
"assets/gold/guide.md": b"allowed",
|
||||
"assets/other/secret.md": b"secret",
|
||||
},
|
||||
reference_roots={"gold": ("assets/gold",)},
|
||||
)
|
||||
|
||||
assert "secret" not in str(exc_info.value)
|
||||
assert "assets/other" not in str(exc_info.value)
|
||||
|
||||
|
||||
def test_empty_reference_files_with_roots_deny_read_and_list_without_io(monkeypatch):
|
||||
"""有根无快照表示根内全部拒绝,read/list 均不得回落磁盘。"""
|
||||
session = CheapSession(
|
||||
game_id="snapshot-empty-files",
|
||||
reference_files={},
|
||||
reference_roots={"gold": ("assets/gold",)},
|
||||
)
|
||||
read_calls = []
|
||||
list_calls = []
|
||||
|
||||
def live_read(path):
|
||||
read_calls.append(path)
|
||||
return {"ok": True, "content": "disk-secret", "truncated": False}
|
||||
|
||||
def live_list(path):
|
||||
list_calls.append(path)
|
||||
return {"ok": True, "entries": ["disk-secret.txt"]}
|
||||
|
||||
monkeypatch.setattr(cheap_toolkit.cheap_run, "read_file", live_read)
|
||||
monkeypatch.setattr(cheap_toolkit.cheap_run, "list_dir", live_list)
|
||||
|
||||
assert _call_tool(session, "read_file", path="assets/gold/unknown.txt").startswith("ERROR:")
|
||||
assert _call_tool(session, "list_dir", path="assets/gold") == ""
|
||||
assert read_calls == []
|
||||
assert list_calls == []
|
||||
|
||||
|
||||
def test_reference_list_path_alias_stays_virtual_without_io(monkeypatch):
|
||||
"""受保护根的规范化别名仍使用虚拟列表,不能触发活目录 I/O。"""
|
||||
session = CheapSession(
|
||||
game_id="snapshot-list-alias",
|
||||
reference_files={"assets/gold/guide.md": b"snapshot"},
|
||||
reference_roots={"gold": ("assets/gold",)},
|
||||
)
|
||||
calls = []
|
||||
|
||||
def live_list(path):
|
||||
calls.append(path)
|
||||
return {"ok": True, "entries": ["disk-secret.txt"]}
|
||||
|
||||
monkeypatch.setattr(cheap_toolkit.cheap_run, "list_dir", live_list)
|
||||
|
||||
assert _call_tool(session, "list_dir", path="assets/gold/./nested/../") == "guide.md"
|
||||
assert calls == []
|
||||
|
||||
|
||||
def test_reference_virtual_list_keeps_toolkit_result_cap():
|
||||
"""虚拟目录投影过大时仍限制在 30KB 工具结果上限内。"""
|
||||
session = CheapSession(
|
||||
game_id="snapshot-list-cap",
|
||||
reference_files={
|
||||
f"assets/gold/file-{index:05d}.txt": b""
|
||||
for index in range(5000)
|
||||
},
|
||||
reference_roots={"gold": ("assets/gold",)},
|
||||
)
|
||||
|
||||
result = _call_tool(session, "list_dir", path="assets/gold")
|
||||
|
||||
assert len(result.encode("utf-8")) == 30_000
|
||||
|
||||
|
||||
def test_cheap_session_keeps_legacy_positional_arguments():
|
||||
"""旧的四个 CheapSession 位置参数仍按原字段接收。"""
|
||||
check = {"ok": True}
|
||||
build = {"ok": True}
|
||||
finished = {"summary": "legacy"}
|
||||
|
||||
session = CheapSession("legacy", check, build, finished)
|
||||
|
||||
assert session.game_id == "legacy"
|
||||
assert session.last_check is check
|
||||
assert session.last_build is build
|
||||
assert session.finished is finished
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
_fns = [v for k, v in sorted(globals().items()) if k.startswith("test_") and callable(v)]
|
||||
_failed = 0
|
||||
|
||||
92
contracts/play-loop/acceptance-provenance-v3.schema.json
Normal file
92
contracts/play-loop/acceptance-provenance-v3.schema.json
Normal file
@ -0,0 +1,92 @@
|
||||
{
|
||||
"$schema": "https://json-schema.org/draft/2020-12/schema",
|
||||
"$id": "https://wanxiang.ai/contracts/play-loop/acceptance-provenance-v3.schema.json",
|
||||
"title": "AcceptanceProvenanceV3",
|
||||
"description": "acceptance-provenance/3。逐字段镜像 acceptance-request/3(含 v3 新增的可选 designRef / referenceAssetRecordIds / consumerRef),并绑定完整 canonical request hash 与被测 artifact。新增可选消费溯源字段 consumedReferenceAssets:本次 acceptance 实际消费的参照资产对账快照,每条为 recordId + role + artifactHash 三元组——对应金标 SoT §7『消费者按 recordId 对账 role + consumerRef + artifactHash 后才能读取资产』的溯源留痕;快照冻结消费时刻的 role 与 hash,注册表后续改动不覆盖历史 provenance。所有 v3 新增字段均可选、不进 required,旧 provenance(v2 字段集)在本 schema 下仍然 valid;v2 的 required 集合与修回 if-then 约束不变。",
|
||||
"type": "object",
|
||||
"required": [
|
||||
"schemaVersion", "gameId", "briefHash", "genre", "templateRoute", "proofProfileId",
|
||||
"proofRegistryVersion", "taskBindingHash", "interactionBinding", "sourceArtifactHash",
|
||||
"parentAcceptanceRequestHash", "repairOrdinal", "acceptanceRequestHash", "artifactHash"
|
||||
],
|
||||
"properties": {
|
||||
"schemaVersion": { "const": "acceptance-provenance/3" },
|
||||
"gameId": { "type": "string", "pattern": "^[A-Za-z0-9][A-Za-z0-9._-]*$" },
|
||||
"briefHash": { "$ref": "#/$defs/sha256" },
|
||||
"genre": { "enum": ["narrative", "trpg", "heritage", "puzzle", "sim-business"] },
|
||||
"templateRoute": { "type": "string", "pattern": "^_template-[a-z0-9-]+$" },
|
||||
"proofProfileId": { "type": "string", "pattern": "^[a-z][a-z0-9-]*\\.[a-z][a-z0-9-]*$" },
|
||||
"proofRegistryVersion": { "type": "string", "minLength": 1 },
|
||||
"taskBindingHash": { "$ref": "#/$defs/sha256" },
|
||||
"interactionBinding": {
|
||||
"anyOf": [
|
||||
{ "$ref": "interaction-binding.schema.json" },
|
||||
{ "type": "null" }
|
||||
]
|
||||
},
|
||||
"sourceArtifactHash": { "anyOf": [{ "$ref": "#/$defs/sha256" }, { "type": "null" }] },
|
||||
"parentAcceptanceRequestHash": { "anyOf": [{ "$ref": "#/$defs/sha256" }, { "type": "null" }] },
|
||||
"repairOrdinal": { "enum": [0, 1] },
|
||||
"acceptanceRequestHash": { "$ref": "#/$defs/sha256" },
|
||||
"artifactHash": { "$ref": "#/$defs/sha256" },
|
||||
"designRef": {
|
||||
"description": "可选,镜像 acceptance-request/3.designRef。",
|
||||
"anyOf": [{ "type": "string", "minLength": 1 }, { "type": "null" }]
|
||||
},
|
||||
"referenceAssetRecordIds": {
|
||||
"description": "可选,镜像 acceptance-request/3.referenceAssetRecordIds(声明消费的 recordId 列表)。",
|
||||
"type": "array",
|
||||
"items": { "$ref": "#/$defs/recordId" }
|
||||
},
|
||||
"consumerRef": {
|
||||
"description": "可选,镜像 acceptance-request/3.consumerRef(消费者身份)。",
|
||||
"anyOf": [{ "type": "string", "minLength": 1 }, { "type": "null" }]
|
||||
},
|
||||
"consumedReferenceAssets": {
|
||||
"description": "可选。本次 acceptance 实际消费的参照资产对账快照列表。每条冻结消费时刻的 recordId + role + artifactHash 三元组,使历史 provenance 不随注册表后续改动漂移;与请求侧 referenceAssetRecordIds 的区别:前者是声明,本字段是实际消费后的核验留痕。",
|
||||
"type": "array",
|
||||
"items": { "$ref": "#/$defs/consumedReferenceAsset" }
|
||||
}
|
||||
},
|
||||
"additionalProperties": false,
|
||||
"allOf": [
|
||||
{
|
||||
"if": { "properties": { "repairOrdinal": { "const": 0 } }, "required": ["repairOrdinal"] },
|
||||
"then": {
|
||||
"properties": {
|
||||
"sourceArtifactHash": { "type": "null" },
|
||||
"parentAcceptanceRequestHash": { "type": "null" }
|
||||
}
|
||||
}
|
||||
},
|
||||
{
|
||||
"if": { "properties": { "repairOrdinal": { "const": 1 } }, "required": ["repairOrdinal"] },
|
||||
"then": {
|
||||
"properties": {
|
||||
"sourceArtifactHash": { "$ref": "#/$defs/sha256" },
|
||||
"parentAcceptanceRequestHash": { "$ref": "#/$defs/sha256" }
|
||||
}
|
||||
}
|
||||
}
|
||||
],
|
||||
"$defs": {
|
||||
"sha256": { "type": "string", "pattern": "^[0-9a-f]{64}$" },
|
||||
"recordId": { "type": "string", "pattern": "^[A-Za-z0-9_][A-Za-z0-9._-]*$" },
|
||||
"consumedReferenceAsset": {
|
||||
"type": "object",
|
||||
"required": ["recordId", "role", "artifactHash"],
|
||||
"properties": {
|
||||
"recordId": { "$ref": "#/$defs/recordId" },
|
||||
"role": {
|
||||
"description": "消费时刻的角色快照,枚举与 ReferenceAssetRecord/1.role 一致。",
|
||||
"enum": ["harness_fixture", "prompt_eval_gold", "generation_exemplar", "game_content_gold"]
|
||||
},
|
||||
"artifactHash": {
|
||||
"description": "消费时刻核验的制品 hash 快照(play-loop 惯例 64 位小写十六进制)。",
|
||||
"$ref": "#/$defs/sha256"
|
||||
}
|
||||
},
|
||||
"additionalProperties": false
|
||||
}
|
||||
}
|
||||
}
|
||||
121
contracts/play-loop/acceptance-provenance-v4.schema.json
Normal file
121
contracts/play-loop/acceptance-provenance-v4.schema.json
Normal file
@ -0,0 +1,121 @@
|
||||
{
|
||||
"$schema": "https://json-schema.org/draft/2020-12/schema",
|
||||
"$id": "https://wanxiang.ai/contracts/play-loop/acceptance-provenance-v4.schema.json",
|
||||
"title": "AcceptanceProvenanceV4",
|
||||
"description": "acceptance-provenance/4。逐字段镜像 acceptance-request/3 与 acceptance-provenance/3 的身份字段,不改写 /3;新增且必填的 referenceAssetVerificationReceipts 以 ref/hash 引用可信消费回执,阻断只有声明没有验证证据的 provenance。",
|
||||
"type": "object",
|
||||
"required": [
|
||||
"schemaVersion", "gameId", "briefHash", "genre", "templateRoute", "proofProfileId",
|
||||
"proofRegistryVersion", "taskBindingHash", "interactionBinding", "sourceArtifactHash",
|
||||
"parentAcceptanceRequestHash", "repairOrdinal", "acceptanceRequestHash", "artifactHash",
|
||||
"referenceAssetVerificationReceipts"
|
||||
],
|
||||
"properties": {
|
||||
"schemaVersion": { "const": "acceptance-provenance/4" },
|
||||
"gameId": { "type": "string", "pattern": "^[A-Za-z0-9][A-Za-z0-9._-]*$" },
|
||||
"briefHash": { "$ref": "#/$defs/sha256" },
|
||||
"genre": { "enum": ["narrative", "trpg", "heritage", "puzzle", "sim-business"] },
|
||||
"templateRoute": { "type": "string", "pattern": "^_template-[a-z0-9-]+$" },
|
||||
"proofProfileId": { "type": "string", "pattern": "^[a-z][a-z0-9-]*\\.[a-z][a-z0-9-]*$" },
|
||||
"proofRegistryVersion": { "type": "string", "minLength": 1 },
|
||||
"taskBindingHash": { "$ref": "#/$defs/sha256" },
|
||||
"interactionBinding": {
|
||||
"anyOf": [
|
||||
{ "$ref": "interaction-binding.schema.json" },
|
||||
{ "type": "null" }
|
||||
]
|
||||
},
|
||||
"sourceArtifactHash": {
|
||||
"anyOf": [{ "$ref": "#/$defs/sha256" }, { "type": "null" }]
|
||||
},
|
||||
"parentAcceptanceRequestHash": {
|
||||
"anyOf": [{ "$ref": "#/$defs/sha256" }, { "type": "null" }]
|
||||
},
|
||||
"repairOrdinal": { "enum": [0, 1] },
|
||||
"acceptanceRequestHash": { "$ref": "#/$defs/sha256" },
|
||||
"artifactHash": { "$ref": "#/$defs/sha256" },
|
||||
"designRef": {
|
||||
"anyOf": [{ "type": "string", "minLength": 1 }, { "type": "null" }]
|
||||
},
|
||||
"referenceAssetRecordIds": {
|
||||
"type": "array",
|
||||
"items": { "$ref": "#/$defs/recordId" }
|
||||
},
|
||||
"consumerRef": {
|
||||
"anyOf": [{ "type": "string", "minLength": 1 }, { "type": "null" }]
|
||||
},
|
||||
"consumedReferenceAssets": {
|
||||
"type": "array",
|
||||
"items": { "$ref": "#/$defs/consumedReferenceAsset" }
|
||||
},
|
||||
"referenceAssetVerificationReceipts": {
|
||||
"type": "array",
|
||||
"minItems": 1,
|
||||
"items": { "$ref": "#/$defs/verificationReceiptRef" }
|
||||
}
|
||||
},
|
||||
"additionalProperties": false,
|
||||
"allOf": [
|
||||
{
|
||||
"if": {
|
||||
"properties": { "repairOrdinal": { "const": 0 } },
|
||||
"required": ["repairOrdinal"]
|
||||
},
|
||||
"then": {
|
||||
"properties": {
|
||||
"sourceArtifactHash": { "type": "null" },
|
||||
"parentAcceptanceRequestHash": { "type": "null" }
|
||||
}
|
||||
}
|
||||
},
|
||||
{
|
||||
"if": {
|
||||
"properties": { "repairOrdinal": { "const": 1 } },
|
||||
"required": ["repairOrdinal"]
|
||||
},
|
||||
"then": {
|
||||
"properties": {
|
||||
"sourceArtifactHash": { "$ref": "#/$defs/sha256" },
|
||||
"parentAcceptanceRequestHash": { "$ref": "#/$defs/sha256" }
|
||||
}
|
||||
}
|
||||
}
|
||||
],
|
||||
"$defs": {
|
||||
"sha256": {
|
||||
"type": "string",
|
||||
"pattern": "^[0-9a-f]{64}$"
|
||||
},
|
||||
"recordId": {
|
||||
"type": "string",
|
||||
"pattern": "^[A-Za-z0-9_][A-Za-z0-9._-]*$"
|
||||
},
|
||||
"relativePath": {
|
||||
"type": "string",
|
||||
"minLength": 1,
|
||||
"maxLength": 1024,
|
||||
"pattern": "^(?!/)(?!.*//)(?!.*(?:^|/)\\.{1,2}(?:/|$))(?!.*\\\\).+$"
|
||||
},
|
||||
"consumedReferenceAsset": {
|
||||
"type": "object",
|
||||
"required": ["recordId", "role", "artifactHash"],
|
||||
"properties": {
|
||||
"recordId": { "$ref": "#/$defs/recordId" },
|
||||
"role": {
|
||||
"enum": ["harness_fixture", "prompt_eval_gold", "generation_exemplar", "game_content_gold"]
|
||||
},
|
||||
"artifactHash": { "$ref": "#/$defs/sha256" }
|
||||
},
|
||||
"additionalProperties": false
|
||||
},
|
||||
"verificationReceiptRef": {
|
||||
"type": "object",
|
||||
"required": ["ref", "hash"],
|
||||
"properties": {
|
||||
"ref": { "$ref": "#/$defs/relativePath" },
|
||||
"hash": { "$ref": "#/$defs/sha256" }
|
||||
},
|
||||
"additionalProperties": false
|
||||
}
|
||||
}
|
||||
}
|
||||
69
contracts/play-loop/acceptance-request-v3.schema.json
Normal file
69
contracts/play-loop/acceptance-request-v3.schema.json
Normal file
@ -0,0 +1,69 @@
|
||||
{
|
||||
"$schema": "https://json-schema.org/draft/2020-12/schema",
|
||||
"$id": "https://wanxiang.ai/contracts/play-loop/acceptance-request-v3.schema.json",
|
||||
"title": "AcceptanceRequestV3",
|
||||
"description": "acceptance-request/3。在 acceptance-request/2 之上为 W-GOLD-LIVE 参照资产消费开口:新增三个可选字段 designRef / referenceAssetRecordIds / consumerRef,全部不进 required,旧 acceptance(不带这三个字段)在本 schema 下仍然 valid——v2 的 required 集合、repairOrdinal 的 if-then 修回约束、interactionBinding 嵌套规则与 canonical acceptanceRequestHash 口径一律不变。交互身份只允许嵌套 InteractionBinding/1 或显式 null。designRef/referenceAssetRecordIds 指向的持久记录定义见 reference-asset-record.schema.json(金标 SoT §7):消费者按 recordId 对账 role + consumerRef + artifactHash 后才能读取资产,未 active 的记录不得被消费为校准锚/生成范例/内容金标。",
|
||||
"type": "object",
|
||||
"required": [
|
||||
"schemaVersion", "gameId", "briefHash", "genre", "templateRoute", "proofProfileId",
|
||||
"proofRegistryVersion", "taskBindingHash", "interactionBinding", "sourceArtifactHash",
|
||||
"parentAcceptanceRequestHash", "repairOrdinal"
|
||||
],
|
||||
"properties": {
|
||||
"schemaVersion": { "const": "acceptance-request/3" },
|
||||
"gameId": { "type": "string", "pattern": "^[A-Za-z0-9][A-Za-z0-9._-]*$" },
|
||||
"briefHash": { "$ref": "#/$defs/sha256" },
|
||||
"genre": { "enum": ["narrative", "trpg", "heritage", "puzzle", "sim-business"] },
|
||||
"templateRoute": { "type": "string", "pattern": "^_template-[a-z0-9-]+$" },
|
||||
"proofProfileId": { "type": "string", "pattern": "^[a-z][a-z0-9-]*\\.[a-z][a-z0-9-]*$" },
|
||||
"proofRegistryVersion": { "type": "string", "minLength": 1 },
|
||||
"taskBindingHash": { "$ref": "#/$defs/sha256" },
|
||||
"interactionBinding": {
|
||||
"anyOf": [
|
||||
{ "$ref": "interaction-binding.schema.json" },
|
||||
{ "type": "null" }
|
||||
]
|
||||
},
|
||||
"sourceArtifactHash": { "anyOf": [{ "$ref": "#/$defs/sha256" }, { "type": "null" }] },
|
||||
"parentAcceptanceRequestHash": { "anyOf": [{ "$ref": "#/$defs/sha256" }, { "type": "null" }] },
|
||||
"repairOrdinal": { "enum": [0, 1] },
|
||||
"designRef": {
|
||||
"description": "可选。本次 acceptance 所依据的已批准 designIntent 引用(持久记录 recordId 或设计文档 ref/路径,如金标 SoT §7.1《山海行纪》追认记录)。不带本字段等价于『未声明设计引用』,与 v2 行为一致。",
|
||||
"anyOf": [{ "type": "string", "minLength": 1 }, { "type": "null" }]
|
||||
},
|
||||
"referenceAssetRecordIds": {
|
||||
"description": "可选。本次 acceptance 声明消费的参照资产 recordId 列表(ReferenceAssetRecord/1.recordId)。仅声明消费意图;实际消费对账快照(role + artifactHash)落在 acceptance-provenance/3 的 consumedReferenceAssets。元素必须形如合法 recordId。",
|
||||
"type": "array",
|
||||
"items": { "$ref": "#/$defs/recordId" }
|
||||
},
|
||||
"consumerRef": {
|
||||
"description": "可选。本次 acceptance 的消费者身份(runner/profile、Actor/Judge/rubric 版本、prompt id@version 或内容评测版本),供参照资产注册表按 recordId 对账消费方。语义与 ReferenceAssetRecord/1.consumerRef 同源。",
|
||||
"anyOf": [{ "type": "string", "minLength": 1 }, { "type": "null" }]
|
||||
}
|
||||
},
|
||||
"additionalProperties": false,
|
||||
"allOf": [
|
||||
{
|
||||
"if": { "properties": { "repairOrdinal": { "const": 0 } }, "required": ["repairOrdinal"] },
|
||||
"then": {
|
||||
"properties": {
|
||||
"sourceArtifactHash": { "type": "null" },
|
||||
"parentAcceptanceRequestHash": { "type": "null" }
|
||||
}
|
||||
}
|
||||
},
|
||||
{
|
||||
"if": { "properties": { "repairOrdinal": { "const": 1 } }, "required": ["repairOrdinal"] },
|
||||
"then": {
|
||||
"properties": {
|
||||
"sourceArtifactHash": { "$ref": "#/$defs/sha256" },
|
||||
"parentAcceptanceRequestHash": { "$ref": "#/$defs/sha256" }
|
||||
}
|
||||
}
|
||||
}
|
||||
],
|
||||
"$defs": {
|
||||
"sha256": { "type": "string", "pattern": "^[0-9a-f]{64}$" },
|
||||
"recordId": { "type": "string", "pattern": "^[A-Za-z0-9_][A-Za-z0-9._-]*$" }
|
||||
}
|
||||
}
|
||||
@ -0,0 +1,294 @@
|
||||
#!/usr/bin/env python3
|
||||
# -*- coding: utf-8 -*-
|
||||
"""生成《山海行纪》参照资产消费清单的确定性命令。"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import argparse
|
||||
import hashlib
|
||||
import json
|
||||
import re
|
||||
import stat
|
||||
import subprocess
|
||||
import sys
|
||||
import unicodedata
|
||||
from pathlib import Path
|
||||
|
||||
|
||||
DEFAULT_REGISTRY = Path("contracts/play-loop/reference-asset-registry.initial.json")
|
||||
DEFAULT_RECORD_ID = "gac-shanhai-xingji"
|
||||
DEFAULT_MANIFEST_ID = "survivor-gold-v1-manifest"
|
||||
MANIFEST_SCHEMA_VERSION = "ReferenceAssetConsumptionManifest/1"
|
||||
CANONICALIZATION = "reference-asset-consumption-manifest/1"
|
||||
DEFAULT_SOURCE_MODE = "worktree"
|
||||
SOURCE_MODES = ("worktree", "clean-archive", "clean_archive")
|
||||
TEMPORARY_OUTPUT_RE = re.compile(r"(^|[._-])(cache|debug|generated|output|tmp|temp)([._-]|$)", re.IGNORECASE)
|
||||
|
||||
|
||||
class ManifestGenerationError(Exception):
|
||||
"""表示清单输入不可信或不完整的稳定生成错误。"""
|
||||
|
||||
def __init__(self, code: str, logical_path: str, detail: str) -> None:
|
||||
super().__init__(detail)
|
||||
self.code = code
|
||||
self.logical_path = logical_path
|
||||
self.detail = detail
|
||||
|
||||
|
||||
def canonical_json_bytes(value: object) -> bytes:
|
||||
"""按项目契约生成无 BOM、无尾随换行的 canonical JSON 字节。"""
|
||||
# sort_keys 使用 Unicode code point 排序,ensure_ascii=False 保留非 ASCII 原字节。
|
||||
return json.dumps(
|
||||
value,
|
||||
ensure_ascii=False,
|
||||
sort_keys=True,
|
||||
separators=(",", ":"),
|
||||
allow_nan=False,
|
||||
).encode("utf-8")
|
||||
|
||||
|
||||
def logical_path(repo_root: Path, path: Path) -> str:
|
||||
"""把路径转换为仓根相对 POSIX NFC 路径,拒绝越出仓根的引用。"""
|
||||
try:
|
||||
relative = path.resolve().relative_to(repo_root.resolve()).as_posix()
|
||||
except ValueError as exc:
|
||||
raise ManifestGenerationError(
|
||||
"reference_path_escape", "<input>", "输入路径必须位于可信仓根内",
|
||||
) from exc
|
||||
normalized = unicodedata.normalize("NFC", relative)
|
||||
if normalized != relative or not relative or relative.startswith("/"):
|
||||
raise ManifestGenerationError("reference_path_invalid", relative or "<empty>", "路径不是 NFC 相对 POSIX 路径")
|
||||
return normalized
|
||||
|
||||
|
||||
def read_regular_file(repo_root: Path, path: Path) -> tuple[str, bytes]:
|
||||
"""以仓内逻辑路径读取普通文件,避免把链接或目录当作消费输入。"""
|
||||
relative = logical_path(repo_root, path)
|
||||
try:
|
||||
file_stat = path.lstat()
|
||||
except FileNotFoundError as exc:
|
||||
raise ManifestGenerationError("reference_missing", relative, "纳入文件不存在") from exc
|
||||
except OSError as exc:
|
||||
raise ManifestGenerationError("reference_unreadable", relative, "纳入文件不可读取") from exc
|
||||
if stat.S_ISLNK(file_stat.st_mode):
|
||||
raise ManifestGenerationError("reference_symlink", relative, "纳入文件不得是 symlink")
|
||||
if not stat.S_ISREG(file_stat.st_mode):
|
||||
raise ManifestGenerationError("reference_not_regular", relative, "纳入对象必须是普通文件")
|
||||
try:
|
||||
raw = path.read_bytes()
|
||||
except OSError as exc:
|
||||
raise ManifestGenerationError("reference_unreadable", relative, "纳入文件不可读取") from exc
|
||||
return relative, raw
|
||||
|
||||
|
||||
def load_registry(repo_root: Path, registry_ref: str, record_id: str) -> dict:
|
||||
"""读取当前 /1 registry,只从 active 记录机械取得已批准 designRef。"""
|
||||
registry_path = repo_root / registry_ref
|
||||
relative, raw = read_regular_file(repo_root, registry_path)
|
||||
try:
|
||||
registry = json.loads(raw.decode("utf-8"))
|
||||
except (UnicodeDecodeError, json.JSONDecodeError) as exc:
|
||||
raise ManifestGenerationError("reference_registry_invalid", relative, "registry 不是 UTF-8 JSON") from exc
|
||||
if not isinstance(registry, dict) or registry.get("schemaVersion") != "ReferenceAssetRegistry/1":
|
||||
raise ManifestGenerationError("reference_registry_invalid", relative, "生成器只接受冻结的 Registry/1 输入")
|
||||
matches = [
|
||||
record for record in registry.get("records", [])
|
||||
if isinstance(record, dict) and record.get("recordId") == record_id
|
||||
]
|
||||
if len(matches) != 1 or matches[0].get("lifecycleStatus") != "active":
|
||||
raise ManifestGenerationError("reference_registry_invalid", relative, "目标记录必须唯一且为 active")
|
||||
design_refs = matches[0].get("designRef")
|
||||
if not isinstance(design_refs, list) or not design_refs or not all(isinstance(ref, str) for ref in design_refs):
|
||||
raise ManifestGenerationError("reference_registry_invalid", relative, "active 记录缺少已批准 designRef")
|
||||
return {"record": matches[0], "designRefs": design_refs}
|
||||
|
||||
|
||||
def normalize_source_mode(mode: str) -> str:
|
||||
"""规范化输入来源模式,默认只允许真实 Git worktree。"""
|
||||
normalized = mode.replace("_", "-")
|
||||
if normalized not in ("worktree", "clean-archive"):
|
||||
raise ManifestGenerationError("reference_source_mode_invalid", "<input>", "输入来源模式不受支持")
|
||||
return normalized
|
||||
|
||||
|
||||
def require_git_worktree(repo_root: Path) -> None:
|
||||
"""确认仓根自身是 Git worktree,禁止把父仓或无元数据归档当作真实仓。"""
|
||||
try:
|
||||
result = subprocess.run(
|
||||
["git", "-C", str(repo_root), "rev-parse", "--show-toplevel"],
|
||||
capture_output=True,
|
||||
text=True,
|
||||
check=False,
|
||||
)
|
||||
except OSError as exc:
|
||||
raise ManifestGenerationError(
|
||||
"reference_git_metadata_unavailable", "<input>", "真实 worktree 模式需要可用的 Git 元数据",
|
||||
) from exc
|
||||
if result.returncode != 0:
|
||||
raise ManifestGenerationError(
|
||||
"reference_git_metadata_unavailable", "<input>", "真实 worktree 模式需要可用的 Git 元数据",
|
||||
)
|
||||
try:
|
||||
git_root = Path(result.stdout.strip()).resolve()
|
||||
except (OSError, ValueError) as exc:
|
||||
raise ManifestGenerationError(
|
||||
"reference_git_metadata_unavailable", "<input>", "真实 worktree 模式需要可用的 Git 元数据",
|
||||
) from exc
|
||||
if not result.stdout.strip() or git_root != repo_root.resolve():
|
||||
raise ManifestGenerationError(
|
||||
"reference_git_metadata_unavailable", "<input>", "真实 worktree 模式需要可用的 Git 元数据",
|
||||
)
|
||||
|
||||
|
||||
def is_git_ignored(repo_root: Path, path: Path) -> bool:
|
||||
"""使用 Git 原生语义判断候选是否 ignored;未知状态一律拒绝继续。"""
|
||||
relative = logical_path(repo_root, path)
|
||||
try:
|
||||
result = subprocess.run(
|
||||
["git", "-C", str(repo_root), "check-ignore", "--quiet", "--", relative],
|
||||
capture_output=True,
|
||||
check=False,
|
||||
)
|
||||
except OSError as exc:
|
||||
raise ManifestGenerationError(
|
||||
"reference_git_check_failed", relative, "Git ignore 检查失败",
|
||||
) from exc
|
||||
if result.returncode == 0:
|
||||
return True
|
||||
if result.returncode == 1:
|
||||
return False
|
||||
raise ManifestGenerationError("reference_git_check_failed", relative, "Git ignore 检查失败")
|
||||
|
||||
|
||||
def is_temporary_output(path: Path) -> bool:
|
||||
"""排除固定白名单下可识别的调试、缓存、生成和临时输出命名。"""
|
||||
return TEMPORARY_OUTPUT_RE.search(path.name) is not None
|
||||
|
||||
|
||||
def selected_paths(
|
||||
repo_root: Path,
|
||||
registry_ref: str,
|
||||
record_id: str,
|
||||
mode: str = DEFAULT_SOURCE_MODE,
|
||||
) -> list[Path]:
|
||||
"""按固定白名单收集输入,再应用 Git 与临时输出排除规则。"""
|
||||
source_mode = normalize_source_mode(mode)
|
||||
if source_mode == "worktree":
|
||||
require_git_worktree(repo_root)
|
||||
game_root = repo_root / "game-runtime/games/shanhai-xingji"
|
||||
fixed = [game_root / name for name in ("README.md", "index.html", "entry.js")]
|
||||
fixed.append(game_root / "assets/manifest.json")
|
||||
fixed.extend(sorted((game_root / "src").glob("*.js"), key=lambda path: logical_path(repo_root, path).encode("utf-8")))
|
||||
fixed.extend(sorted((game_root / "assets/atlas").glob("*.json"), key=lambda path: logical_path(repo_root, path).encode("utf-8")))
|
||||
design_refs = load_registry(repo_root, registry_ref, record_id)["designRefs"]
|
||||
fixed.extend(repo_root / ref for ref in design_refs)
|
||||
unique: dict[str, Path] = {}
|
||||
for path in fixed:
|
||||
relative = logical_path(repo_root, path)
|
||||
if relative in unique:
|
||||
raise ManifestGenerationError("reference_path_invalid", relative, "消费路径不得重复")
|
||||
unique[relative] = path
|
||||
selected = []
|
||||
for relative in sorted(unique, key=lambda value: value.encode("utf-8")):
|
||||
path = unique[relative]
|
||||
if is_temporary_output(path):
|
||||
continue
|
||||
if source_mode == "worktree" and is_git_ignored(repo_root, path):
|
||||
continue
|
||||
selected.append(path)
|
||||
return selected
|
||||
|
||||
|
||||
def build_manifest(
|
||||
repo_root: Path,
|
||||
registry_ref: str = DEFAULT_REGISTRY.as_posix(),
|
||||
record_id: str = DEFAULT_RECORD_ID,
|
||||
manifest_id: str = DEFAULT_MANIFEST_ID,
|
||||
mode: str = DEFAULT_SOURCE_MODE,
|
||||
) -> dict:
|
||||
"""读取批准范围并生成只含原始 size/hash 的 manifest 对象。"""
|
||||
entries = []
|
||||
for path in selected_paths(repo_root, registry_ref, record_id, mode):
|
||||
relative, raw = read_regular_file(repo_root, path)
|
||||
entries.append({
|
||||
"path": relative,
|
||||
"size": len(raw),
|
||||
"sha256": hashlib.sha256(raw).hexdigest(),
|
||||
})
|
||||
# 路径排序使用 UTF-8 字节,确保不同语言实现得到同一数组顺序。
|
||||
entries.sort(key=lambda entry: entry["path"].encode("utf-8"))
|
||||
return {
|
||||
"schemaVersion": MANIFEST_SCHEMA_VERSION,
|
||||
"manifestId": manifest_id,
|
||||
"canonicalization": CANONICALIZATION,
|
||||
"entries": entries,
|
||||
}
|
||||
|
||||
|
||||
def parse_args(argv: list[str]) -> argparse.Namespace:
|
||||
"""解析生成命令参数;默认值固定到仓内首个 active 金标。"""
|
||||
parser = argparse.ArgumentParser(description="生成参照资产消费 manifest")
|
||||
parser.add_argument("--repo-root", type=Path, required=True, help="可信仓根")
|
||||
parser.add_argument("--output", type=Path, required=True, help="manifest 输出文件")
|
||||
parser.add_argument("--registry", default=DEFAULT_REGISTRY.as_posix(), help="仓根相对 Registry/1 路径")
|
||||
parser.add_argument("--record-id", default=DEFAULT_RECORD_ID, help="active 记录 ID")
|
||||
parser.add_argument("--manifest-id", default=DEFAULT_MANIFEST_ID, help="manifest 逻辑 ID")
|
||||
parser.add_argument(
|
||||
"--mode",
|
||||
"--source-mode",
|
||||
dest="mode",
|
||||
choices=SOURCE_MODES,
|
||||
default=DEFAULT_SOURCE_MODE,
|
||||
help="输入来源模式;默认 worktree,归档必须显式选择 clean-archive",
|
||||
)
|
||||
parser.add_argument(
|
||||
"--clean-archive",
|
||||
dest="mode",
|
||||
action="store_const",
|
||||
const="clean-archive",
|
||||
help="显式选择无 Git 元数据的 clean-archive 输入",
|
||||
)
|
||||
return parser.parse_args(argv)
|
||||
|
||||
|
||||
def main(argv: list[str] | None = None) -> int:
|
||||
"""执行生成并输出可审计的稳定摘要,不输出文件内容或机器绝对路径。"""
|
||||
args = parse_args(argv or sys.argv[1:])
|
||||
repo_root = args.repo_root.resolve()
|
||||
try:
|
||||
manifest = build_manifest(
|
||||
repo_root,
|
||||
registry_ref=args.registry,
|
||||
record_id=args.record_id,
|
||||
manifest_id=args.manifest_id,
|
||||
mode=args.mode,
|
||||
)
|
||||
raw = canonical_json_bytes(manifest)
|
||||
args.output.parent.mkdir(parents=True, exist_ok=True)
|
||||
args.output.write_bytes(raw)
|
||||
output_ref = args.output.resolve().relative_to(repo_root).as_posix() if args.output.resolve().is_relative_to(repo_root) else "<output>"
|
||||
print(
|
||||
f"generated reference manifest recordId={args.record_id} "
|
||||
f"entries={len(manifest['entries'])} sha256={hashlib.sha256(raw).hexdigest()} "
|
||||
f"path={output_ref}",
|
||||
)
|
||||
return 0
|
||||
except ManifestGenerationError as exc:
|
||||
print(
|
||||
f"reference_manifest_generation_failed code={exc.code} "
|
||||
f"recordId={args.record_id} path={exc.logical_path}",
|
||||
file=sys.stderr,
|
||||
)
|
||||
return 1
|
||||
except (OSError, ValueError, TypeError) as exc:
|
||||
# 未预期的本地 I/O/参数错误仍只报告稳定类别,不泄露绝对路径或输入内容。
|
||||
print(
|
||||
f"reference_manifest_generation_failed code=reference_generation_error "
|
||||
f"recordId={args.record_id} path=<input> detail={type(exc).__name__}",
|
||||
file=sys.stderr,
|
||||
)
|
||||
return 1
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
raise SystemExit(main())
|
||||
102
contracts/play-loop/historical-11-expectations.json
Normal file
102
contracts/play-loop/historical-11-expectations.json
Normal file
@ -0,0 +1,102 @@
|
||||
{
|
||||
"schemaVersion": "historical-expectations/1",
|
||||
"topic": "cheap-gen-acceptance-v3-historical-11",
|
||||
"canonicalSource": "docs/agent-specs/2026-07-13-生成线验收v3-可信证据闭环-设计.md §5.2",
|
||||
"description": "历史 11 局固定预期表(机器可读 fixture)。v3 切换前的历史重放闸门据此逐局对账:3 个 narrative 正例应保留(expect=accept);5 个旧假阳在没有新硬证时不得 accept(expect=not_accept,accept 即确认假阳放行);puzzle-r2 真代码 bug 应 reject(expect=reject);2 个疑似假阴必须先经真人真浏览器定标(expect=needs_human,定标前闸门不得自动 accept、不计真好通过率)。gid 沿用波 3 复核命名(见设计档 §5.2 与父 plan 失效模式三案)。",
|
||||
"gatePolicy": {
|
||||
"falsePositiveReleaseMax": 0,
|
||||
"perGameExpectationMatchRequired": true,
|
||||
"needsHumanAutoAcceptAllowed": false,
|
||||
"testerErrorIsMismatch": true
|
||||
},
|
||||
"expectations": [
|
||||
{
|
||||
"gid": "narrative-r1",
|
||||
"genre": "narrative",
|
||||
"expected": "accept",
|
||||
"category": "positive_control",
|
||||
"critical": false,
|
||||
"rationale": "narrative 正例应保留:v2 时代加落点回显后 12 步真玩通到结局卡,v3 必须仍能 accept,否则是验收器回退"
|
||||
},
|
||||
{
|
||||
"gid": "narrative-r2",
|
||||
"genre": "narrative",
|
||||
"expected": "accept",
|
||||
"category": "positive_control",
|
||||
"critical": false,
|
||||
"rationale": "narrative 正例应保留:有后果分支选择 + 明确结局的硬证齐全,v3 必须 accept"
|
||||
},
|
||||
{
|
||||
"gid": "narrative-r3",
|
||||
"genre": "narrative",
|
||||
"expected": "accept",
|
||||
"category": "positive_control",
|
||||
"critical": false,
|
||||
"rationale": "narrative 正例应保留:结局与已走分支一致的有序事件链可复算,v3 必须 accept"
|
||||
},
|
||||
{
|
||||
"gid": "trpg-r1",
|
||||
"genre": "trpg",
|
||||
"expected": "not_accept",
|
||||
"category": "known_false_positive",
|
||||
"critical": true,
|
||||
"rationale": "旧假阳:v2 自证循环下被接受(战斗局卡在升级页仍 accept)。没有新硬证时不得 accept;reject 或 inconclusive 均可"
|
||||
},
|
||||
{
|
||||
"gid": "puzzle-r1",
|
||||
"genre": "puzzle",
|
||||
"expected": "not_accept",
|
||||
"category": "known_false_positive",
|
||||
"critical": true,
|
||||
"rationale": "旧假阳:解谜局未归位仍被接受。没有新硬证(真实归位事件链 + 终盘截图)时不得 accept"
|
||||
},
|
||||
{
|
||||
"gid": "trpg-r2",
|
||||
"genre": "trpg",
|
||||
"expected": "not_accept",
|
||||
"category": "known_false_positive",
|
||||
"critical": true,
|
||||
"rationale": "旧假阳:局部动画/插件日志被当成输入有效。没有新硬证时不得 accept"
|
||||
},
|
||||
{
|
||||
"gid": "heritage-r2",
|
||||
"genre": "heritage",
|
||||
"expected": "not_accept",
|
||||
"category": "known_false_positive",
|
||||
"critical": true,
|
||||
"rationale": "旧假阳:非遗局零成品仍被接受。没有新硬证(成品画面 + 工序有序系列)时不得 accept"
|
||||
},
|
||||
{
|
||||
"gid": "sim-business-r2",
|
||||
"genre": "sim-business",
|
||||
"expected": "not_accept",
|
||||
"category": "known_false_positive",
|
||||
"critical": true,
|
||||
"rationale": "旧假阳:经营局零订单仍被接受(失效模式三案之一)。没有新硬证(order-served 事件链)时不得 accept"
|
||||
},
|
||||
{
|
||||
"gid": "puzzle-r2",
|
||||
"genre": "puzzle",
|
||||
"expected": "reject",
|
||||
"category": "true_bug",
|
||||
"critical": true,
|
||||
"rationale": "真代码 bug:硬证已证明可复现的产物缺陷,应 reject(accept 即假阳放行;inconclusive 与固定预期不符,需查验收器为何拿不到硬证)"
|
||||
},
|
||||
{
|
||||
"gid": "heritage-r1",
|
||||
"genre": "heritage",
|
||||
"expected": "needs_human",
|
||||
"category": "suspected_false_negative",
|
||||
"critical": false,
|
||||
"rationale": "疑似假阴:必须先经真人真浏览器定标,未定标前不进入真好通过率;闸门对 accept 判失败(定标前不得自动 accept),reject/inconclusive 挂起等创始人定标,不自动放行也不自动通过"
|
||||
},
|
||||
{
|
||||
"gid": "sim-business-r1",
|
||||
"genre": "sim-business",
|
||||
"expected": "needs_human",
|
||||
"category": "suspected_false_negative",
|
||||
"critical": false,
|
||||
"rationale": "疑似假阴:节奏窗口在模型推理期间耗尽可能误拒好游戏,必须先真人真浏览器定标;定标前闸门不得自动 accept,其余结果挂起等创始人定标"
|
||||
}
|
||||
]
|
||||
}
|
||||
@ -0,0 +1,48 @@
|
||||
{
|
||||
"$schema": "https://json-schema.org/draft/2020-12/schema",
|
||||
"$id": "https://wanxiang.ai/contracts/play-loop/reference-asset-consumption-manifest.schema.json",
|
||||
"title": "ReferenceAssetConsumptionManifestV1",
|
||||
"description": "ReferenceAssetConsumptionManifest/1。manifest 使用项目自有 reference-asset-consumption-manifest/1 canonical JSON 字节口径;entries 必须按 path 升序、路径必须为 NFC 的仓内相对 POSIX 路径,size 为非负 JSON 安全整数,sha256 为小写 SHA-256。",
|
||||
"type": "object",
|
||||
"required": ["schemaVersion", "manifestId", "canonicalization", "entries"],
|
||||
"properties": {
|
||||
"schemaVersion": { "const": "ReferenceAssetConsumptionManifest/1" },
|
||||
"manifestId": {
|
||||
"type": "string",
|
||||
"pattern": "^[A-Za-z0-9_][A-Za-z0-9._-]*$"
|
||||
},
|
||||
"canonicalization": { "const": "reference-asset-consumption-manifest/1" },
|
||||
"entries": {
|
||||
"type": "array",
|
||||
"minItems": 1,
|
||||
"items": { "$ref": "#/$defs/entry" }
|
||||
}
|
||||
},
|
||||
"additionalProperties": false,
|
||||
"$defs": {
|
||||
"entry": {
|
||||
"type": "object",
|
||||
"required": ["path", "size", "sha256"],
|
||||
"properties": {
|
||||
"path": { "$ref": "#/$defs/relativePath" },
|
||||
"size": {
|
||||
"type": "integer",
|
||||
"minimum": 0,
|
||||
"maximum": 9007199254740991
|
||||
},
|
||||
"sha256": { "$ref": "#/$defs/sha256" }
|
||||
},
|
||||
"additionalProperties": false
|
||||
},
|
||||
"sha256": {
|
||||
"type": "string",
|
||||
"pattern": "^[0-9a-f]{64}$"
|
||||
},
|
||||
"relativePath": {
|
||||
"type": "string",
|
||||
"minLength": 1,
|
||||
"maxLength": 1024,
|
||||
"pattern": "^(?!/)(?!.*//)(?!.*(?:^|/)\\.{1,2}(?:/|$))(?!.*\\\\).+$"
|
||||
}
|
||||
}
|
||||
}
|
||||
@ -0,0 +1,10 @@
|
||||
{
|
||||
"schemaVersion": "ReferenceAssetConsumptionPolicy/1",
|
||||
"policyId": "survivor-gold-v1",
|
||||
"recordId": "gac-shanhai-xingji",
|
||||
"role": "game_content_gold",
|
||||
"consumerRef": "generation-runtime@reference-assets/2",
|
||||
"route": "survivor-gold",
|
||||
"autoSelect": false,
|
||||
"mode": "frozen_preflight"
|
||||
}
|
||||
@ -0,0 +1,21 @@
|
||||
{
|
||||
"$schema": "https://json-schema.org/draft/2020-12/schema",
|
||||
"$id": "https://wanxiang.ai/contracts/play-loop/reference-asset-consumption-policy.schema.json",
|
||||
"title": "ReferenceAssetConsumptionPolicyV1",
|
||||
"description": "ReferenceAssetConsumptionPolicy/1。survivor-gold-v1 是唯一冻结 preflight 策略:只能绑定已签认 game_content_gold,禁止自动选择,消费模式固定为 frozen_preflight。",
|
||||
"type": "object",
|
||||
"required": [
|
||||
"schemaVersion", "policyId", "recordId", "role", "consumerRef", "route", "autoSelect", "mode"
|
||||
],
|
||||
"properties": {
|
||||
"schemaVersion": { "const": "ReferenceAssetConsumptionPolicy/1" },
|
||||
"policyId": { "const": "survivor-gold-v1" },
|
||||
"recordId": { "const": "gac-shanhai-xingji" },
|
||||
"role": { "const": "game_content_gold" },
|
||||
"consumerRef": { "const": "generation-runtime@reference-assets/2" },
|
||||
"route": { "const": "survivor-gold" },
|
||||
"autoSelect": { "const": false },
|
||||
"mode": { "const": "frozen_preflight" }
|
||||
},
|
||||
"additionalProperties": false
|
||||
}
|
||||
159
contracts/play-loop/reference-asset-record-v2.schema.json
Normal file
159
contracts/play-loop/reference-asset-record-v2.schema.json
Normal file
@ -0,0 +1,159 @@
|
||||
{
|
||||
"$schema": "https://json-schema.org/draft/2020-12/schema",
|
||||
"$id": "https://wanxiang.ai/contracts/play-loop/reference-asset-record-v2.schema.json",
|
||||
"title": "ReferenceAssetRecordV2",
|
||||
"description": "ReferenceAssetRecord/2。冻结并扩展 ReferenceAssetRecord/1,不修改 /1 版本线。candidate/migration_pending 可以保留三项新身份为 null;active 必须具备运行制品、消费清单和清单 hash;retired 若保留历史签认则必须完整保留身份字段。",
|
||||
"type": "object",
|
||||
"required": [
|
||||
"schemaVersion", "recordId", "role", "lifecycleStatus",
|
||||
"assetRef", "assetVersion", "artifactHash", "evidenceRefs"
|
||||
],
|
||||
"properties": {
|
||||
"schemaVersion": { "const": "ReferenceAssetRecord/2" },
|
||||
"recordId": {
|
||||
"type": "string",
|
||||
"pattern": "^[A-Za-z0-9_][A-Za-z0-9._-]*$"
|
||||
},
|
||||
"role": {
|
||||
"enum": ["harness_fixture", "prompt_eval_gold", "generation_exemplar", "game_content_gold"]
|
||||
},
|
||||
"lifecycleStatus": {
|
||||
"enum": ["candidate", "migration_pending", "active", "retired"]
|
||||
},
|
||||
"assetRef": { "$ref": "#/$defs/relativePath" },
|
||||
"assetVersion": { "type": "string", "minLength": 1 },
|
||||
"artifactHash": { "$ref": "#/$defs/sha256" },
|
||||
"consumerRef": {
|
||||
"anyOf": [{ "type": "string", "minLength": 1 }, { "type": "null" }]
|
||||
},
|
||||
"designRef": {
|
||||
"anyOf": [
|
||||
{
|
||||
"type": "array",
|
||||
"items": { "type": "string", "minLength": 1 },
|
||||
"minItems": 1
|
||||
},
|
||||
{ "type": "null" }
|
||||
]
|
||||
},
|
||||
"evidenceRefs": {
|
||||
"type": "array",
|
||||
"items": { "type": "string", "minLength": 1 }
|
||||
},
|
||||
"signedBy": {
|
||||
"anyOf": [{ "type": "string", "minLength": 1 }, { "type": "null" }]
|
||||
},
|
||||
"signedAt": {
|
||||
"anyOf": [
|
||||
{
|
||||
"type": "string",
|
||||
"pattern": "^\\d{4}-\\d{2}-\\d{2}(T\\d{2}:\\d{2}(:\\d{2})?(Z|[+-]\\d{2}:?\\d{2})?)?$"
|
||||
},
|
||||
{ "type": "null" }
|
||||
]
|
||||
},
|
||||
"artifactRef": {
|
||||
"anyOf": [{ "$ref": "#/$defs/relativePath" }, { "type": "null" }]
|
||||
},
|
||||
"consumptionManifestRef": {
|
||||
"anyOf": [{ "$ref": "#/$defs/relativePath" }, { "type": "null" }]
|
||||
},
|
||||
"consumptionManifestHash": {
|
||||
"anyOf": [{ "$ref": "#/$defs/sha256" }, { "type": "null" }]
|
||||
}
|
||||
},
|
||||
"additionalProperties": false,
|
||||
"allOf": [
|
||||
{
|
||||
"if": {
|
||||
"properties": { "lifecycleStatus": { "const": "active" } },
|
||||
"required": ["lifecycleStatus"]
|
||||
},
|
||||
"then": {
|
||||
"required": [
|
||||
"consumerRef", "signedBy", "signedAt",
|
||||
"artifactRef", "consumptionManifestRef", "consumptionManifestHash"
|
||||
],
|
||||
"properties": {
|
||||
"consumerRef": { "type": "string", "minLength": 1 },
|
||||
"signedBy": { "type": "string", "minLength": 1 },
|
||||
"signedAt": {
|
||||
"type": "string",
|
||||
"pattern": "^\\d{4}-\\d{2}-\\d{2}(T\\d{2}:\\d{2}(:\\d{2})?(Z|[+-]\\d{2}:?\\d{2})?)?$"
|
||||
},
|
||||
"artifactRef": { "$ref": "#/$defs/relativePath" },
|
||||
"consumptionManifestRef": { "$ref": "#/$defs/relativePath" },
|
||||
"consumptionManifestHash": { "$ref": "#/$defs/sha256" }
|
||||
}
|
||||
}
|
||||
},
|
||||
{
|
||||
"if": {
|
||||
"properties": {
|
||||
"role": { "const": "game_content_gold" },
|
||||
"lifecycleStatus": { "const": "active" }
|
||||
},
|
||||
"required": ["role", "lifecycleStatus"]
|
||||
},
|
||||
"then": {
|
||||
"properties": {
|
||||
"designRef": {
|
||||
"type": "array",
|
||||
"items": { "type": "string", "minLength": 1 },
|
||||
"minItems": 1
|
||||
}
|
||||
},
|
||||
"required": ["designRef"]
|
||||
}
|
||||
},
|
||||
{
|
||||
"if": {
|
||||
"properties": { "lifecycleStatus": { "const": "retired" } },
|
||||
"required": ["lifecycleStatus"],
|
||||
"anyOf": [
|
||||
{
|
||||
"required": ["consumerRef"],
|
||||
"properties": { "consumerRef": { "type": "string", "minLength": 1 } }
|
||||
},
|
||||
{
|
||||
"required": ["signedBy"],
|
||||
"properties": { "signedBy": { "type": "string", "minLength": 1 } }
|
||||
},
|
||||
{
|
||||
"required": ["signedAt"],
|
||||
"properties": { "signedAt": { "type": "string", "minLength": 1 } }
|
||||
}
|
||||
]
|
||||
},
|
||||
"then": {
|
||||
"required": [
|
||||
"consumerRef", "signedBy", "signedAt",
|
||||
"artifactRef", "consumptionManifestRef", "consumptionManifestHash"
|
||||
],
|
||||
"properties": {
|
||||
"consumerRef": { "type": "string", "minLength": 1 },
|
||||
"signedBy": { "type": "string", "minLength": 1 },
|
||||
"signedAt": {
|
||||
"type": "string",
|
||||
"pattern": "^\\d{4}-\\d{2}-\\d{2}(T\\d{2}:\\d{2}(:\\d{2})?(Z|[+-]\\d{2}:?\\d{2})?)?$"
|
||||
},
|
||||
"artifactRef": { "$ref": "#/$defs/relativePath" },
|
||||
"consumptionManifestRef": { "$ref": "#/$defs/relativePath" },
|
||||
"consumptionManifestHash": { "$ref": "#/$defs/sha256" }
|
||||
}
|
||||
}
|
||||
}
|
||||
],
|
||||
"$defs": {
|
||||
"sha256": {
|
||||
"type": "string",
|
||||
"pattern": "^[0-9a-f]{64}$"
|
||||
},
|
||||
"relativePath": {
|
||||
"type": "string",
|
||||
"minLength": 1,
|
||||
"maxLength": 1024,
|
||||
"pattern": "^(?!/)(?!.*//)(?!.*(?:^|/)\\.{1,2}(?:/|$))(?!.*\\\\).+$"
|
||||
}
|
||||
}
|
||||
}
|
||||
109
contracts/play-loop/reference-asset-record.schema.json
Normal file
109
contracts/play-loop/reference-asset-record.schema.json
Normal file
@ -0,0 +1,109 @@
|
||||
{
|
||||
"$schema": "https://json-schema.org/draft/2020-12/schema",
|
||||
"$id": "https://wanxiang.ai/contracts/play-loop/reference-asset-record.schema.json",
|
||||
"title": "ReferenceAssetRecordV1",
|
||||
"description": "ReferenceAssetRecord/1。可复用校准与参照资产的持久消费记录,字段定义权威源为金标 SoT docs/architecture/产品/游戏内容金标.md §7。分类以消费记录为互斥单位、不以物理目录为单位;消费者只接收 recordId,再从持久注册表读取并核验完整记录。role 与 lifecycleStatus 正交:candidate/migration_pending 不是第五种角色;同一物理资产新增第二个角色必须新建第二条记录,禁止就地改写 role。active 态必须补齐 consumerRef/signedBy/signedAt(SoT『active 时必填』);game_content_gold 升 active 还必须有已批准 designIntent(SoT §7 字段表 designRef 行)。candidate/migration_pending 处于补证阶段,允许这些字段暂为 null,但 artifactHash 等制品身份在建记录时即须补齐——迁移清单阶段的占位 hash 须在注册表 migrationNotes 中显式标注,active 前必须替换为真 hash。",
|
||||
"type": "object",
|
||||
"required": [
|
||||
"schemaVersion", "recordId", "role", "lifecycleStatus",
|
||||
"assetRef", "assetVersion", "artifactHash", "evidenceRefs"
|
||||
],
|
||||
"properties": {
|
||||
"schemaVersion": { "const": "ReferenceAssetRecord/1" },
|
||||
"recordId": {
|
||||
"description": "稳定唯一的记录标识,创建后不可改;所有消费日志与证据引用它。允许下划线开头以兼容 _fewshot-*/_template-* 既有资产名。",
|
||||
"type": "string",
|
||||
"pattern": "^[A-Za-z0-9_][A-Za-z0-9._-]*$"
|
||||
},
|
||||
"role": {
|
||||
"description": "消费角色,四选一且互斥。harness_fixture=校工具(Actor/Judge/专项夹具);prompt_eval_gold=校裁判(rubric 正反例与校准卷);generation_exemplar=引导生成器(模板与范例);game_content_gold=完整内容参照(须九维人工签认,夹具不得冒充)。",
|
||||
"enum": ["harness_fixture", "prompt_eval_gold", "generation_exemplar", "game_content_gold"]
|
||||
},
|
||||
"lifecycleStatus": {
|
||||
"description": "生命周期状态,与 role 正交。candidate=候选待补证;migration_pending=迁移清单已登记、制品身份待冻结绑定;active=责任人核对设计与证据后签认、可供 live 消费;retired=退休只读。候选与迁移状态不是第五种角色。",
|
||||
"enum": ["candidate", "migration_pending", "active", "retired"]
|
||||
},
|
||||
"assetRef": {
|
||||
"description": "指向唯一冻结制品的引用(仓内路径或稳定标识)。hash 漂移必须新建版本或记录,不得覆盖旧证据。migration_pending 阶段可为占位引用,active 前必须绑定真制品。",
|
||||
"type": "string",
|
||||
"minLength": 1
|
||||
},
|
||||
"assetVersion": {
|
||||
"description": "制品版本(如夹具 r3、组件 1.0.0)。迁移时冻结;未冻结资产以占位值登记并在注册表 migrationNotes 标注。",
|
||||
"type": "string",
|
||||
"minLength": 1
|
||||
},
|
||||
"artifactHash": {
|
||||
"description": "冻结制品的 SHA-256(play-loop 惯例 64 位小写十六进制、无算法前缀)。迁移清单阶段允许占位 hash(须标注),active 前必须替换为真 hash;不得以覆盖旧 hash 的方式『更新』证据。",
|
||||
"$ref": "#/$defs/sha256"
|
||||
},
|
||||
"consumerRef": {
|
||||
"description": "具体消费者身份:runner/profile、Actor/Judge/rubric 版本、prompt id@version 或内容评测版本。active 时必填(见 allOf 条件);非 active 态为 null 表示尚未接线消费。",
|
||||
"anyOf": [{ "type": "string", "minLength": 1 }, { "type": "null" }]
|
||||
},
|
||||
"designRef": {
|
||||
"description": "已批准 designIntent 的引用列表(设计文档仓内路径或 recordId/ref)。game_content_gold 升 active 必须指向已批准 designIntent(见 allOf 条件);完整游戏来源的 generation_exemplar 同理;其它角色不适用时置 null,理由写入注册表 migrationNotes。candidate 阶段允许已批 designIntent 先行登记(如《山海行纪》2026-07-23 契约追认)。",
|
||||
"anyOf": [
|
||||
{ "type": "array", "items": { "type": "string", "minLength": 1 }, "minItems": 1 },
|
||||
{ "type": "null" }
|
||||
]
|
||||
},
|
||||
"evidenceRefs": {
|
||||
"description": "角色对应的校准、真玩、审计或内容证据引用;不得只写自然语言结论。建记录时必须存在本字段;暂无证据引用时以空数组显式表达『待补』,不允许省略字段。",
|
||||
"type": "array",
|
||||
"items": { "type": "string", "minLength": 1 }
|
||||
},
|
||||
"signedBy": {
|
||||
"description": "角色责任人签认身份。active 时必填;内容金标须满足金标 SoT §6 人工终审。candidate/migration_pending 为 null。",
|
||||
"anyOf": [{ "type": "string", "minLength": 1 }, { "type": "null" }]
|
||||
},
|
||||
"signedAt": {
|
||||
"description": "签认时间(ISO 8601 日期或日期时间)。active 时必填;candidate/migration_pending 为 null。",
|
||||
"anyOf": [
|
||||
{
|
||||
"type": "string",
|
||||
"pattern": "^\\d{4}-\\d{2}-\\d{2}(T\\d{2}:\\d{2}(:\\d{2})?(Z|[+-]\\d{2}:?\\d{2})?)?$"
|
||||
},
|
||||
{ "type": "null" }
|
||||
]
|
||||
}
|
||||
},
|
||||
"additionalProperties": false,
|
||||
"allOf": [
|
||||
{
|
||||
"if": {
|
||||
"properties": { "lifecycleStatus": { "const": "active" } },
|
||||
"required": ["lifecycleStatus"]
|
||||
},
|
||||
"then": {
|
||||
"properties": {
|
||||
"consumerRef": { "type": "string", "minLength": 1 },
|
||||
"signedBy": { "type": "string", "minLength": 1 },
|
||||
"signedAt": {
|
||||
"type": "string",
|
||||
"pattern": "^\\d{4}-\\d{2}-\\d{2}(T\\d{2}:\\d{2}(:\\d{2})?(Z|[+-]\\d{2}:?\\d{2})?)?$"
|
||||
}
|
||||
},
|
||||
"required": ["consumerRef", "signedBy", "signedAt"]
|
||||
}
|
||||
},
|
||||
{
|
||||
"if": {
|
||||
"properties": {
|
||||
"role": { "const": "game_content_gold" },
|
||||
"lifecycleStatus": { "const": "active" }
|
||||
},
|
||||
"required": ["role", "lifecycleStatus"]
|
||||
},
|
||||
"then": {
|
||||
"properties": {
|
||||
"designRef": { "type": "array", "items": { "type": "string", "minLength": 1 }, "minItems": 1 }
|
||||
},
|
||||
"required": ["designRef"]
|
||||
}
|
||||
}
|
||||
],
|
||||
"$defs": {
|
||||
"sha256": { "type": "string", "pattern": "^[0-9a-f]{64}$" }
|
||||
}
|
||||
}
|
||||
25
contracts/play-loop/reference-asset-registry-v2.schema.json
Normal file
25
contracts/play-loop/reference-asset-registry-v2.schema.json
Normal file
@ -0,0 +1,25 @@
|
||||
{
|
||||
"$schema": "https://json-schema.org/draft/2020-12/schema",
|
||||
"$id": "https://wanxiang.ai/contracts/play-loop/reference-asset-registry-v2.schema.json",
|
||||
"title": "ReferenceAssetRegistryV2",
|
||||
"description": "ReferenceAssetRegistry/2。只接收 ReferenceAssetRecord/2;recordId 的跨条目唯一性、migration-list 的生命周期约束和注释闭包由 validate.py 语义层强制。",
|
||||
"type": "object",
|
||||
"required": ["schemaVersion", "registryVersion", "sourceOfTruth", "records"],
|
||||
"properties": {
|
||||
"schemaVersion": { "const": "ReferenceAssetRegistry/2" },
|
||||
"registryVersion": { "type": "string", "minLength": 1 },
|
||||
"sourceOfTruth": { "type": "string", "minLength": 1 },
|
||||
"records": {
|
||||
"type": "array",
|
||||
"items": { "$ref": "reference-asset-record-v2.schema.json" }
|
||||
},
|
||||
"migrationNotes": {
|
||||
"type": "object",
|
||||
"patternProperties": {
|
||||
"^[A-Za-z0-9_][A-Za-z0-9._-]*$": { "type": "string", "minLength": 1 }
|
||||
},
|
||||
"additionalProperties": false
|
||||
}
|
||||
},
|
||||
"additionalProperties": false
|
||||
}
|
||||
244
contracts/play-loop/reference-asset-registry.initial.json
Normal file
244
contracts/play-loop/reference-asset-registry.initial.json
Normal file
@ -0,0 +1,244 @@
|
||||
{
|
||||
"schemaVersion": "ReferenceAssetRegistry/1",
|
||||
"registryVersion": "2026-07-27.first-active-gold",
|
||||
"sourceOfTruth": "docs/architecture/产品/游戏内容金标.md §7",
|
||||
"records": [
|
||||
{
|
||||
"schemaVersion": "ReferenceAssetRecord/1",
|
||||
"recordId": "gold-m3-gem-r3",
|
||||
"role": "harness_fixture",
|
||||
"lifecycleStatus": "migration_pending",
|
||||
"assetRef": "game-runtime/games/_wg1-gen/gold-m3-gem-r3",
|
||||
"assetVersion": "r3",
|
||||
"artifactHash": "94075fb645952bd068c8429042a0e82247ffe66e72c48951e1af68603d71f4fc",
|
||||
"consumerRef": null,
|
||||
"designRef": null,
|
||||
"evidenceRefs": [],
|
||||
"signedBy": null,
|
||||
"signedAt": null
|
||||
},
|
||||
{
|
||||
"schemaVersion": "ReferenceAssetRecord/1",
|
||||
"recordId": "gold-m3-candy-r3",
|
||||
"role": "harness_fixture",
|
||||
"lifecycleStatus": "migration_pending",
|
||||
"assetRef": "game-runtime/games/_wg1-gen/gold-m3-candy-r3",
|
||||
"assetVersion": "r3",
|
||||
"artifactHash": "e83de56f53ba50cf20a489377aadc900d9f2dd3cb1f772bea8b05c14486228f7",
|
||||
"consumerRef": null,
|
||||
"designRef": null,
|
||||
"evidenceRefs": [],
|
||||
"signedBy": null,
|
||||
"signedAt": null
|
||||
},
|
||||
{
|
||||
"schemaVersion": "ReferenceAssetRecord/1",
|
||||
"recordId": "gold-m3-fruit-r3",
|
||||
"role": "harness_fixture",
|
||||
"lifecycleStatus": "migration_pending",
|
||||
"assetRef": "game-runtime/games/_wg1-gen/gold-m3-fruit-r3",
|
||||
"assetVersion": "r3",
|
||||
"artifactHash": "ea14ba864ab483aa953864cb6cd376f77a606124141a354988f086a2b049aabc",
|
||||
"consumerRef": null,
|
||||
"designRef": null,
|
||||
"evidenceRefs": [],
|
||||
"signedBy": null,
|
||||
"signedAt": null
|
||||
},
|
||||
{
|
||||
"schemaVersion": "ReferenceAssetRecord/1",
|
||||
"recordId": "gold-m3-porcelain-r3",
|
||||
"role": "harness_fixture",
|
||||
"lifecycleStatus": "migration_pending",
|
||||
"assetRef": "game-runtime/games/_wg1-gen/gold-m3-porcelain-r3",
|
||||
"assetVersion": "r3",
|
||||
"artifactHash": "7065c844fbb358a7375409a35e524950e2d82b6e4bb18ebcc94fc3f35924ebbb",
|
||||
"consumerRef": null,
|
||||
"designRef": null,
|
||||
"evidenceRefs": [],
|
||||
"signedBy": null,
|
||||
"signedAt": null
|
||||
},
|
||||
{
|
||||
"schemaVersion": "ReferenceAssetRecord/1",
|
||||
"recordId": "gold-m3-rune-r3",
|
||||
"role": "harness_fixture",
|
||||
"lifecycleStatus": "migration_pending",
|
||||
"assetRef": "game-runtime/games/_wg1-gen/gold-m3-rune-r3",
|
||||
"assetVersion": "r3",
|
||||
"artifactHash": "a0a9192241bbd222611cef76bb7f2fec0c149f726d1fb4e8d29e9555befba54e",
|
||||
"consumerRef": null,
|
||||
"designRef": null,
|
||||
"evidenceRefs": [],
|
||||
"signedBy": null,
|
||||
"signedAt": null
|
||||
},
|
||||
{
|
||||
"schemaVersion": "ReferenceAssetRecord/1",
|
||||
"recordId": "_fewshot-feiyi",
|
||||
"role": "generation_exemplar",
|
||||
"lifecycleStatus": "migration_pending",
|
||||
"assetRef": "game-runtime/games/_fewshot-feiyi",
|
||||
"assetVersion": "unfrozen-2026-07-25",
|
||||
"artifactHash": "5a706ed24767dae43630e030c91e9a1958052429fd59062b2ceeac244353aac3",
|
||||
"consumerRef": null,
|
||||
"designRef": null,
|
||||
"evidenceRefs": [],
|
||||
"signedBy": null,
|
||||
"signedAt": null
|
||||
},
|
||||
{
|
||||
"schemaVersion": "ReferenceAssetRecord/1",
|
||||
"recordId": "_fewshot-puzzle",
|
||||
"role": "generation_exemplar",
|
||||
"lifecycleStatus": "migration_pending",
|
||||
"assetRef": "game-runtime/games/_fewshot-puzzle",
|
||||
"assetVersion": "unfrozen-2026-07-25",
|
||||
"artifactHash": "88cca39a7029ff81584c55df9498c7cd0f50cda3121561628fb6bd88f80ab63c",
|
||||
"consumerRef": null,
|
||||
"designRef": null,
|
||||
"evidenceRefs": [],
|
||||
"signedBy": null,
|
||||
"signedAt": null
|
||||
},
|
||||
{
|
||||
"schemaVersion": "ReferenceAssetRecord/1",
|
||||
"recordId": "_template-feiyi",
|
||||
"role": "generation_exemplar",
|
||||
"lifecycleStatus": "migration_pending",
|
||||
"assetRef": "game-runtime/games/_template-feiyi",
|
||||
"assetVersion": "unfrozen-2026-07-25",
|
||||
"artifactHash": "58c2694aa814127a6773693506566c6afb351004ca5f42b5d409d7dd628f6faf",
|
||||
"consumerRef": null,
|
||||
"designRef": null,
|
||||
"evidenceRefs": [],
|
||||
"signedBy": null,
|
||||
"signedAt": null
|
||||
},
|
||||
{
|
||||
"schemaVersion": "ReferenceAssetRecord/1",
|
||||
"recordId": "_template-puzzle",
|
||||
"role": "generation_exemplar",
|
||||
"lifecycleStatus": "migration_pending",
|
||||
"assetRef": "game-runtime/games/_template-puzzle",
|
||||
"assetVersion": "unfrozen-2026-07-25",
|
||||
"artifactHash": "9ceee42869560d5806184544ff9668bc8c56a4b582a74dc6bb8fab9511281fa9",
|
||||
"consumerRef": null,
|
||||
"designRef": null,
|
||||
"evidenceRefs": [],
|
||||
"signedBy": null,
|
||||
"signedAt": null
|
||||
},
|
||||
{
|
||||
"schemaVersion": "ReferenceAssetRecord/1",
|
||||
"recordId": "_template-shop",
|
||||
"role": "generation_exemplar",
|
||||
"lifecycleStatus": "migration_pending",
|
||||
"assetRef": "game-runtime/games/_template-shop",
|
||||
"assetVersion": "unfrozen-2026-07-25",
|
||||
"artifactHash": "565ede6017f57bba4652e3910cec45df06f5789773816039d1ed891fb9518e2f",
|
||||
"consumerRef": null,
|
||||
"designRef": null,
|
||||
"evidenceRefs": [],
|
||||
"signedBy": null,
|
||||
"signedAt": null
|
||||
},
|
||||
{
|
||||
"schemaVersion": "ReferenceAssetRecord/1",
|
||||
"recordId": "_template-story",
|
||||
"role": "generation_exemplar",
|
||||
"lifecycleStatus": "migration_pending",
|
||||
"assetRef": "game-runtime/games/_template-story",
|
||||
"assetVersion": "unfrozen-2026-07-25",
|
||||
"artifactHash": "c8f2814e5c535bf0b1c12e9c363f78eca555e83fd5860c4fcfd5a0438ba1006a",
|
||||
"consumerRef": null,
|
||||
"designRef": null,
|
||||
"evidenceRefs": [],
|
||||
"signedBy": null,
|
||||
"signedAt": null
|
||||
},
|
||||
{
|
||||
"schemaVersion": "ReferenceAssetRecord/1",
|
||||
"recordId": "_template-trpg",
|
||||
"role": "generation_exemplar",
|
||||
"lifecycleStatus": "migration_pending",
|
||||
"assetRef": "game-runtime/games/_template-trpg",
|
||||
"assetVersion": "unfrozen-2026-07-25",
|
||||
"artifactHash": "67d19d4970f8e378de6d59370fc6e0a0aa49aebdd3b3237f64835793235e74dc",
|
||||
"consumerRef": null,
|
||||
"designRef": null,
|
||||
"evidenceRefs": [],
|
||||
"signedBy": null,
|
||||
"signedAt": null
|
||||
},
|
||||
{
|
||||
"schemaVersion": "ReferenceAssetRecord/1",
|
||||
"recordId": "gac-shanhai-xingji",
|
||||
"role": "game_content_gold",
|
||||
"lifecycleStatus": "active",
|
||||
"assetRef": "game-runtime/games/shanhai-xingji",
|
||||
"assetVersion": "map1-vertical-slice-r1",
|
||||
"artifactHash": "1c760811ec435fe0f3b5ba79aa8c4fcc44119e019b3e1240ce2c51e55e25870b",
|
||||
"consumerRef": "generation-runtime@reference-assets/1",
|
||||
"designRef": [
|
||||
"docs/agent-specs/2026-07-06-北极星顶级线-肉鸽割草-开发设计书.md",
|
||||
"docs/agent-specs/2026-07-06-山海宇宙设定与美术音频管线-选型材料.md"
|
||||
],
|
||||
"evidenceRefs": [
|
||||
"game-runtime/games/shanhai-xingji/evidence/realization-evidence-status.md",
|
||||
"game-runtime/games/shanhai-xingji/evidence/round-14/00-win-result.json",
|
||||
"game-runtime/games/shanhai-xingji/evidence/round-15/00-round15-summary.json",
|
||||
"game-runtime/games/shanhai-xingji/evidence/round-17-gold-lock/qa-report.md",
|
||||
"game-runtime/games/shanhai-xingji/evidence/round-17-gold-lock/browser-evidence.json"
|
||||
],
|
||||
"signedBy": "创始人",
|
||||
"signedAt": "2026-07-27T13:30:50Z"
|
||||
},
|
||||
{
|
||||
"schemaVersion": "ReferenceAssetRecord/1",
|
||||
"recordId": "gac-shanhai-xunyi-lu",
|
||||
"role": "game_content_gold",
|
||||
"lifecycleStatus": "candidate",
|
||||
"assetRef": "pending-binding-fable-shanhai-xunyi-lu",
|
||||
"assetVersion": "unfrozen-2026-07-25",
|
||||
"artifactHash": "17dc264e89ed8541fde97029cb87a7426af69a694cfdc70b87c19da0bf0b6f33",
|
||||
"consumerRef": null,
|
||||
"designRef": null,
|
||||
"evidenceRefs": [],
|
||||
"signedBy": null,
|
||||
"signedAt": null
|
||||
},
|
||||
{
|
||||
"schemaVersion": "ReferenceAssetRecord/1",
|
||||
"recordId": "gac-yeshi-yitiaojie",
|
||||
"role": "game_content_gold",
|
||||
"lifecycleStatus": "candidate",
|
||||
"assetRef": "pending-binding-fable-yeshi-yitiaojie",
|
||||
"assetVersion": "unfrozen-2026-07-25",
|
||||
"artifactHash": "4d6ce34b6ee8374c2d04f2b933978e457638e630684d4c993281a422824487e7",
|
||||
"consumerRef": null,
|
||||
"designRef": null,
|
||||
"evidenceRefs": [],
|
||||
"signedBy": null,
|
||||
"signedAt": null
|
||||
}
|
||||
],
|
||||
"migrationNotes": {
|
||||
"_registry": "本文件由金标 SoT §7 约束。2026-07-27 起《山海行纪》地图1完整20分钟纵切版为首款 active game_content_gold;其余条目仍是迁移清单。SoT 表第 194/195 行(Match-3 确定性 fixture 如 attempt-019、_shared Node 回归 → harness_fixture;各品类 rubric 正反例与 Actor/Judge 校准卷 → prompt_eval_gold)因具体制品 ID 未逐个敲定,本快照未枚举,待 ID 敲定后补登记。",
|
||||
"gold-m3-gem-r3": "artifactHash 为 migration_pending 占位(sha256('placeholder:gold-m3-gem-r3')),active 前必须替换为整树冻结真 hash;历史 ID 保留作证据引用不改,新产物改用 m3-cal-*;consumerRef 迁移时绑定 runner/profile;evidenceRefs 迁移时绑定专项校准证据。",
|
||||
"gold-m3-candy-r3": "artifactHash 为 migration_pending 占位(sha256('placeholder:gold-m3-candy-r3')),active 前必须替换为整树冻结真 hash;consumerRef/evidenceRefs 迁移时绑定。",
|
||||
"gold-m3-fruit-r3": "artifactHash 为 migration_pending 占位(sha256('placeholder:gold-m3-fruit-r3')),active 前必须替换为整树冻结真 hash;consumerRef/evidenceRefs 迁移时绑定。",
|
||||
"gold-m3-porcelain-r3": "artifactHash 为 migration_pending 占位(sha256('placeholder:gold-m3-porcelain-r3')),active 前必须替换为整树冻结真 hash;consumerRef/evidenceRefs 迁移时绑定。",
|
||||
"gold-m3-rune-r3": "artifactHash 为 migration_pending 占位(sha256('placeholder:gold-m3-rune-r3')),active 前必须替换为整树冻结真 hash;consumerRef/evidenceRefs 迁移时绑定。",
|
||||
"_fewshot-feiyi": "artifactHash 为 migration_pending 占位,assetVersion 未冻结;迁移时冻结版本/hash、来源、适用品类与消费它的 prompt 版本;完成前不得新增 live 消费;模板/范例非完整游戏,designRef 不适用(非完整游戏来源的 generation_exemplar)。",
|
||||
"_fewshot-puzzle": "artifactHash 为 migration_pending 占位,assetVersion 未冻结;迁移时冻结版本/hash、来源、适用品类与消费它的 prompt 版本;完成前不得新增 live 消费。",
|
||||
"_template-feiyi": "artifactHash 为 migration_pending 占位,assetVersion 未冻结;迁移时冻结版本/hash、来源、适用品类与消费它的 prompt 版本;完成前不得新增 live 消费。",
|
||||
"_template-puzzle": "artifactHash 为 migration_pending 占位,assetVersion 未冻结;迁移时冻结版本/hash、来源、适用品类与消费它的 prompt 版本;完成前不得新增 live 消费。",
|
||||
"_template-shop": "artifactHash 为 migration_pending 占位,assetVersion 未冻结;迁移时冻结版本/hash、来源、适用品类与消费它的 prompt 版本;完成前不得新增 live 消费。",
|
||||
"_template-story": "artifactHash 为 migration_pending 占位,assetVersion 未冻结;迁移时冻结版本/hash、来源、适用品类与消费它的 prompt 版本;完成前不得新增 live 消费。",
|
||||
"_template-trpg": "artifactHash 为 migration_pending 占位,assetVersion 未冻结;迁移时冻结版本/hash、来源、适用品类与消费它的 prompt 版本;完成前不得新增 live 消费。",
|
||||
"gac-shanhai-xingji": "active:创始人于 2026-07-27 正式签认地图1《裂谷原》完整20分钟纵切版,音频批准范围为 SFX-only。artifactHash 是同日从当前源重新构建 dist/shanhai-bundle.js 后计算的真实 SHA-256;地图2–5、29兽正式素材与6首BGM属于后续完整版,不在本记录承诺范围。consumerRef 对应 cheap_studio.py → cheap_verify.build_v3_reference_asset_generation_constraints 的生成运行时参照资产注入门。",
|
||||
"gac-shanhai-xunyi-lu": "candidate:designIntent 尚未追认(designRef 为 null);assetRef/artifactHash 为 candidate 占位(资产树待定位绑定);签认前没有正式 game_content_gold 记录。",
|
||||
"gac-yeshi-yitiaojie": "candidate:designIntent 尚未追认(designRef 为 null);assetRef/artifactHash 为 candidate 占位(资产树待定位绑定);签认前没有正式 game_content_gold 记录。"
|
||||
}
|
||||
}
|
||||
33
contracts/play-loop/reference-asset-registry.schema.json
Normal file
33
contracts/play-loop/reference-asset-registry.schema.json
Normal file
@ -0,0 +1,33 @@
|
||||
{
|
||||
"$schema": "https://json-schema.org/draft/2020-12/schema",
|
||||
"$id": "https://wanxiang.ai/contracts/play-loop/reference-asset-registry.schema.json",
|
||||
"title": "ReferenceAssetRegistryV1",
|
||||
"description": "ReferenceAssetRegistry/1。ReferenceAssetRecord/1 的持久注册表/索引结构,权威源为金标 SoT docs/architecture/产品/游戏内容金标.md §7。records 内 recordId 必须全局唯一(JSON Schema 无法表达跨元素字段唯一性,由 validate.py 语义层 _semantic_validate_reference_asset_registry 强制)。当前阶段(registryVersion 带 migration-list 后缀)只是金标 SoT 迁移清单的机器可读编码,不是可供新 live 消费的机器注册表:消费者接线与 formal registry 封存(registryHash 等)由 W-GOLD-LIVE 后续检查点落地,届时 schema 升版引入;任何新增或扩大消费都必须等对应持久记录变成 active。sourceOfTruth 固定指向金标 SoT,防止注册表与标准定义漂移。",
|
||||
"type": "object",
|
||||
"required": ["schemaVersion", "registryVersion", "sourceOfTruth", "records"],
|
||||
"properties": {
|
||||
"schemaVersion": { "const": "ReferenceAssetRegistry/1" },
|
||||
"registryVersion": {
|
||||
"description": "注册表快照版本。迁移清单阶段形如 <日期>.migration-list;正式机器注册表阶段另起版本线。",
|
||||
"type": "string",
|
||||
"minLength": 1
|
||||
},
|
||||
"sourceOfTruth": {
|
||||
"description": "标准定义 SoT 引用(金标 §7),注册表条目与之冲突时以 SoT 为准。",
|
||||
"type": "string",
|
||||
"minLength": 1
|
||||
},
|
||||
"records": {
|
||||
"description": "ReferenceAssetRecord/1 记录列表;recordId 唯一性由语义层强制。",
|
||||
"type": "array",
|
||||
"items": { "$ref": "reference-asset-record.schema.json" }
|
||||
},
|
||||
"migrationNotes": {
|
||||
"description": "迁移清单阶段注释表:key 为 recordId,value 为该记录的占位项说明与迁移要求(如占位 hash/assetRef 待绑定的具体内容)。仅服务迁移登记透明度,不构成记录模型字段;正式 active 消费以记录本身为准。",
|
||||
"type": "object",
|
||||
"propertyNames": { "pattern": "^[A-Za-z0-9_][A-Za-z0-9._-]*$" },
|
||||
"additionalProperties": { "type": "string", "minLength": 1 }
|
||||
}
|
||||
},
|
||||
"additionalProperties": false
|
||||
}
|
||||
289
contracts/play-loop/reference-asset-registry.v2.initial.json
Normal file
289
contracts/play-loop/reference-asset-registry.v2.initial.json
Normal file
@ -0,0 +1,289 @@
|
||||
{
|
||||
"schemaVersion": "ReferenceAssetRegistry/2",
|
||||
"registryVersion": "2026-07-27.trusted-release-v1",
|
||||
"sourceOfTruth": "docs/architecture/产品/游戏内容金标.md §7",
|
||||
"records": [
|
||||
{
|
||||
"schemaVersion": "ReferenceAssetRecord/2",
|
||||
"recordId": "gold-m3-gem-r3",
|
||||
"role": "harness_fixture",
|
||||
"lifecycleStatus": "migration_pending",
|
||||
"assetRef": "game-runtime/games/_wg1-gen/gold-m3-gem-r3",
|
||||
"assetVersion": "r3",
|
||||
"artifactHash": "94075fb645952bd068c8429042a0e82247ffe66e72c48951e1af68603d71f4fc",
|
||||
"consumerRef": null,
|
||||
"designRef": null,
|
||||
"evidenceRefs": [],
|
||||
"signedBy": null,
|
||||
"signedAt": null,
|
||||
"artifactRef": null,
|
||||
"consumptionManifestRef": null,
|
||||
"consumptionManifestHash": null
|
||||
},
|
||||
{
|
||||
"schemaVersion": "ReferenceAssetRecord/2",
|
||||
"recordId": "gold-m3-candy-r3",
|
||||
"role": "harness_fixture",
|
||||
"lifecycleStatus": "migration_pending",
|
||||
"assetRef": "game-runtime/games/_wg1-gen/gold-m3-candy-r3",
|
||||
"assetVersion": "r3",
|
||||
"artifactHash": "e83de56f53ba50cf20a489377aadc900d9f2dd3cb1f772bea8b05c14486228f7",
|
||||
"consumerRef": null,
|
||||
"designRef": null,
|
||||
"evidenceRefs": [],
|
||||
"signedBy": null,
|
||||
"signedAt": null,
|
||||
"artifactRef": null,
|
||||
"consumptionManifestRef": null,
|
||||
"consumptionManifestHash": null
|
||||
},
|
||||
{
|
||||
"schemaVersion": "ReferenceAssetRecord/2",
|
||||
"recordId": "gold-m3-fruit-r3",
|
||||
"role": "harness_fixture",
|
||||
"lifecycleStatus": "migration_pending",
|
||||
"assetRef": "game-runtime/games/_wg1-gen/gold-m3-fruit-r3",
|
||||
"assetVersion": "r3",
|
||||
"artifactHash": "ea14ba864ab483aa953864cb6cd376f77a606124141a354988f086a2b049aabc",
|
||||
"consumerRef": null,
|
||||
"designRef": null,
|
||||
"evidenceRefs": [],
|
||||
"signedBy": null,
|
||||
"signedAt": null,
|
||||
"artifactRef": null,
|
||||
"consumptionManifestRef": null,
|
||||
"consumptionManifestHash": null
|
||||
},
|
||||
{
|
||||
"schemaVersion": "ReferenceAssetRecord/2",
|
||||
"recordId": "gold-m3-porcelain-r3",
|
||||
"role": "harness_fixture",
|
||||
"lifecycleStatus": "migration_pending",
|
||||
"assetRef": "game-runtime/games/_wg1-gen/gold-m3-porcelain-r3",
|
||||
"assetVersion": "r3",
|
||||
"artifactHash": "7065c844fbb358a7375409a35e524950e2d82b6e4bb18ebcc94fc3f35924ebbb",
|
||||
"consumerRef": null,
|
||||
"designRef": null,
|
||||
"evidenceRefs": [],
|
||||
"signedBy": null,
|
||||
"signedAt": null,
|
||||
"artifactRef": null,
|
||||
"consumptionManifestRef": null,
|
||||
"consumptionManifestHash": null
|
||||
},
|
||||
{
|
||||
"schemaVersion": "ReferenceAssetRecord/2",
|
||||
"recordId": "gold-m3-rune-r3",
|
||||
"role": "harness_fixture",
|
||||
"lifecycleStatus": "migration_pending",
|
||||
"assetRef": "game-runtime/games/_wg1-gen/gold-m3-rune-r3",
|
||||
"assetVersion": "r3",
|
||||
"artifactHash": "a0a9192241bbd222611cef76bb7f2fec0c149f726d1fb4e8d29e9555befba54e",
|
||||
"consumerRef": null,
|
||||
"designRef": null,
|
||||
"evidenceRefs": [],
|
||||
"signedBy": null,
|
||||
"signedAt": null,
|
||||
"artifactRef": null,
|
||||
"consumptionManifestRef": null,
|
||||
"consumptionManifestHash": null
|
||||
},
|
||||
{
|
||||
"schemaVersion": "ReferenceAssetRecord/2",
|
||||
"recordId": "_fewshot-feiyi",
|
||||
"role": "generation_exemplar",
|
||||
"lifecycleStatus": "migration_pending",
|
||||
"assetRef": "game-runtime/games/_fewshot-feiyi",
|
||||
"assetVersion": "unfrozen-2026-07-25",
|
||||
"artifactHash": "5a706ed24767dae43630e030c91e9a1958052429fd59062b2ceeac244353aac3",
|
||||
"consumerRef": null,
|
||||
"designRef": null,
|
||||
"evidenceRefs": [],
|
||||
"signedBy": null,
|
||||
"signedAt": null,
|
||||
"artifactRef": null,
|
||||
"consumptionManifestRef": null,
|
||||
"consumptionManifestHash": null
|
||||
},
|
||||
{
|
||||
"schemaVersion": "ReferenceAssetRecord/2",
|
||||
"recordId": "_fewshot-puzzle",
|
||||
"role": "generation_exemplar",
|
||||
"lifecycleStatus": "migration_pending",
|
||||
"assetRef": "game-runtime/games/_fewshot-puzzle",
|
||||
"assetVersion": "unfrozen-2026-07-25",
|
||||
"artifactHash": "88cca39a7029ff81584c55df9498c7cd0f50cda3121561628fb6bd88f80ab63c",
|
||||
"consumerRef": null,
|
||||
"designRef": null,
|
||||
"evidenceRefs": [],
|
||||
"signedBy": null,
|
||||
"signedAt": null,
|
||||
"artifactRef": null,
|
||||
"consumptionManifestRef": null,
|
||||
"consumptionManifestHash": null
|
||||
},
|
||||
{
|
||||
"schemaVersion": "ReferenceAssetRecord/2",
|
||||
"recordId": "_template-feiyi",
|
||||
"role": "generation_exemplar",
|
||||
"lifecycleStatus": "migration_pending",
|
||||
"assetRef": "game-runtime/games/_template-feiyi",
|
||||
"assetVersion": "unfrozen-2026-07-25",
|
||||
"artifactHash": "58c2694aa814127a6773693506566c6afb351004ca5f42b5d409d7dd628f6faf",
|
||||
"consumerRef": null,
|
||||
"designRef": null,
|
||||
"evidenceRefs": [],
|
||||
"signedBy": null,
|
||||
"signedAt": null,
|
||||
"artifactRef": null,
|
||||
"consumptionManifestRef": null,
|
||||
"consumptionManifestHash": null
|
||||
},
|
||||
{
|
||||
"schemaVersion": "ReferenceAssetRecord/2",
|
||||
"recordId": "_template-puzzle",
|
||||
"role": "generation_exemplar",
|
||||
"lifecycleStatus": "migration_pending",
|
||||
"assetRef": "game-runtime/games/_template-puzzle",
|
||||
"assetVersion": "unfrozen-2026-07-25",
|
||||
"artifactHash": "9ceee42869560d5806184544ff9668bc8c56a4b582a74dc6bb8fab9511281fa9",
|
||||
"consumerRef": null,
|
||||
"designRef": null,
|
||||
"evidenceRefs": [],
|
||||
"signedBy": null,
|
||||
"signedAt": null,
|
||||
"artifactRef": null,
|
||||
"consumptionManifestRef": null,
|
||||
"consumptionManifestHash": null
|
||||
},
|
||||
{
|
||||
"schemaVersion": "ReferenceAssetRecord/2",
|
||||
"recordId": "_template-shop",
|
||||
"role": "generation_exemplar",
|
||||
"lifecycleStatus": "migration_pending",
|
||||
"assetRef": "game-runtime/games/_template-shop",
|
||||
"assetVersion": "unfrozen-2026-07-25",
|
||||
"artifactHash": "565ede6017f57bba4652e3910cec45df06f5789773816039d1ed891fb9518e2f",
|
||||
"consumerRef": null,
|
||||
"designRef": null,
|
||||
"evidenceRefs": [],
|
||||
"signedBy": null,
|
||||
"signedAt": null,
|
||||
"artifactRef": null,
|
||||
"consumptionManifestRef": null,
|
||||
"consumptionManifestHash": null
|
||||
},
|
||||
{
|
||||
"schemaVersion": "ReferenceAssetRecord/2",
|
||||
"recordId": "_template-story",
|
||||
"role": "generation_exemplar",
|
||||
"lifecycleStatus": "migration_pending",
|
||||
"assetRef": "game-runtime/games/_template-story",
|
||||
"assetVersion": "unfrozen-2026-07-25",
|
||||
"artifactHash": "c8f2814e5c535bf0b1c12e9c363f78eca555e83fd5860c4fcfd5a0438ba1006a",
|
||||
"consumerRef": null,
|
||||
"designRef": null,
|
||||
"evidenceRefs": [],
|
||||
"signedBy": null,
|
||||
"signedAt": null,
|
||||
"artifactRef": null,
|
||||
"consumptionManifestRef": null,
|
||||
"consumptionManifestHash": null
|
||||
},
|
||||
{
|
||||
"schemaVersion": "ReferenceAssetRecord/2",
|
||||
"recordId": "_template-trpg",
|
||||
"role": "generation_exemplar",
|
||||
"lifecycleStatus": "migration_pending",
|
||||
"assetRef": "game-runtime/games/_template-trpg",
|
||||
"assetVersion": "unfrozen-2026-07-25",
|
||||
"artifactHash": "67d19d4970f8e378de6d59370fc6e0a0aa49aebdd3b3237f64835793235e74dc",
|
||||
"consumerRef": null,
|
||||
"designRef": null,
|
||||
"evidenceRefs": [],
|
||||
"signedBy": null,
|
||||
"signedAt": null,
|
||||
"artifactRef": null,
|
||||
"consumptionManifestRef": null,
|
||||
"consumptionManifestHash": null
|
||||
},
|
||||
{
|
||||
"schemaVersion": "ReferenceAssetRecord/2",
|
||||
"recordId": "gac-shanhai-xingji",
|
||||
"role": "game_content_gold",
|
||||
"lifecycleStatus": "active",
|
||||
"assetRef": "game-runtime/games/shanhai-xingji",
|
||||
"assetVersion": "map1-vertical-slice-r1",
|
||||
"artifactHash": "1c760811ec435fe0f3b5ba79aa8c4fcc44119e019b3e1240ce2c51e55e25870b",
|
||||
"consumerRef": "generation-runtime@reference-assets/2",
|
||||
"designRef": [
|
||||
"docs/agent-specs/2026-07-06-北极星顶级线-肉鸽割草-开发设计书.md",
|
||||
"docs/agent-specs/2026-07-06-山海宇宙设定与美术音频管线-选型材料.md"
|
||||
],
|
||||
"evidenceRefs": [
|
||||
"game-runtime/games/shanhai-xingji/evidence/realization-evidence-status.md",
|
||||
"game-runtime/games/shanhai-xingji/evidence/round-14/00-win-result.json",
|
||||
"game-runtime/games/shanhai-xingji/evidence/round-15/00-round15-summary.json",
|
||||
"game-runtime/games/shanhai-xingji/evidence/round-17-gold-lock/qa-report.md",
|
||||
"game-runtime/games/shanhai-xingji/evidence/round-17-gold-lock/browser-evidence.json"
|
||||
],
|
||||
"signedBy": "创始人",
|
||||
"signedAt": "2026-07-27T13:30:50Z",
|
||||
"artifactRef": "game-runtime/games/shanhai-xingji/dist/shanhai-bundle.js",
|
||||
"consumptionManifestRef": "game-runtime/games/shanhai-xingji/reference/consumption-manifest.json",
|
||||
"consumptionManifestHash": "0f600598cfb842b213f7d54ae1be1d07fe277a3e66b3e4ab3ad21b73dea320d9"
|
||||
},
|
||||
{
|
||||
"schemaVersion": "ReferenceAssetRecord/2",
|
||||
"recordId": "gac-shanhai-xunyi-lu",
|
||||
"role": "game_content_gold",
|
||||
"lifecycleStatus": "candidate",
|
||||
"assetRef": "pending-binding-fable-shanhai-xunyi-lu",
|
||||
"assetVersion": "unfrozen-2026-07-25",
|
||||
"artifactHash": "17dc264e89ed8541fde97029cb87a7426af69a694cfdc70b87c19da0bf0b6f33",
|
||||
"consumerRef": null,
|
||||
"designRef": null,
|
||||
"evidenceRefs": [],
|
||||
"signedBy": null,
|
||||
"signedAt": null,
|
||||
"artifactRef": null,
|
||||
"consumptionManifestRef": null,
|
||||
"consumptionManifestHash": null
|
||||
},
|
||||
{
|
||||
"schemaVersion": "ReferenceAssetRecord/2",
|
||||
"recordId": "gac-yeshi-yitiaojie",
|
||||
"role": "game_content_gold",
|
||||
"lifecycleStatus": "candidate",
|
||||
"assetRef": "pending-binding-fable-yeshi-yitiaojie",
|
||||
"assetVersion": "unfrozen-2026-07-25",
|
||||
"artifactHash": "4d6ce34b6ee8374c2d04f2b933978e457638e630684d4c993281a422824487e7",
|
||||
"consumerRef": null,
|
||||
"designRef": null,
|
||||
"evidenceRefs": [],
|
||||
"signedBy": null,
|
||||
"signedAt": null,
|
||||
"artifactRef": null,
|
||||
"consumptionManifestRef": null,
|
||||
"consumptionManifestHash": null
|
||||
}
|
||||
],
|
||||
"migrationNotes": {
|
||||
"_registry": "本文件由金标 SoT §7 约束。2026-07-27 起《山海行纪》地图1完整20分钟纵切版为首款 active game_content_gold;其余条目仍是迁移清单。SoT 表第 194/195 行(Match-3 确定性 fixture 如 attempt-019、_shared Node 回归 → harness_fixture;各品类 rubric 正反例与 Actor/Judge 校准卷 → prompt_eval_gold)因具体制品 ID 未逐个敲定,本快照未枚举,待 ID 敲定后补登记。",
|
||||
"gold-m3-gem-r3": "artifactHash 为 migration_pending 占位(sha256('placeholder:gold-m3-gem-r3')),active 前必须替换为整树冻结真 hash;历史 ID 保留作证据引用不改,新产物改用 m3-cal-*;consumerRef 迁移时绑定 runner/profile;evidenceRefs 迁移时绑定专项校准证据。",
|
||||
"gold-m3-candy-r3": "artifactHash 为 migration_pending 占位(sha256('placeholder:gold-m3-candy-r3')),active 前必须替换为整树冻结真 hash;consumerRef/evidenceRefs 迁移时绑定。",
|
||||
"gold-m3-fruit-r3": "artifactHash 为 migration_pending 占位(sha256('placeholder:gold-m3-fruit-r3')),active 前必须替换为整树冻结真 hash;consumerRef/evidenceRefs 迁移时绑定。",
|
||||
"gold-m3-porcelain-r3": "artifactHash 为 migration_pending 占位(sha256('placeholder:gold-m3-porcelain-r3')),active 前必须替换为整树冻结真 hash;consumerRef/evidenceRefs 迁移时绑定。",
|
||||
"gold-m3-rune-r3": "artifactHash 为 migration_pending 占位(sha256('placeholder:gold-m3-rune-r3')),active 前必须替换为整树冻结真 hash;consumerRef/evidenceRefs 迁移时绑定。",
|
||||
"_fewshot-feiyi": "artifactHash 为 migration_pending 占位,assetVersion 未冻结;迁移时冻结版本/hash、来源、适用品类与消费它的 prompt 版本;完成前不得新增 live 消费;模板/范例非完整游戏,designRef 不适用(非完整游戏来源的 generation_exemplar)。",
|
||||
"_fewshot-puzzle": "artifactHash 为 migration_pending 占位,assetVersion 未冻结;迁移时冻结版本/hash、来源、适用品类与消费它的 prompt 版本;完成前不得新增 live 消费。",
|
||||
"_template-feiyi": "artifactHash 为 migration_pending 占位,assetVersion 未冻结;迁移时冻结版本/hash、来源、适用品类与消费它的 prompt 版本;完成前不得新增 live 消费。",
|
||||
"_template-puzzle": "artifactHash 为 migration_pending 占位,assetVersion 未冻结;迁移时冻结版本/hash、来源、适用品类与消费它的 prompt 版本;完成前不得新增 live 消费。",
|
||||
"_template-shop": "artifactHash 为 migration_pending 占位,assetVersion 未冻结;迁移时冻结版本/hash、来源、适用品类与消费它的 prompt 版本;完成前不得新增 live 消费。",
|
||||
"_template-story": "artifactHash 为 migration_pending 占位,assetVersion 未冻结;迁移时冻结版本/hash、来源、适用品类与消费它的 prompt 版本;完成前不得新增 live 消费。",
|
||||
"_template-trpg": "artifactHash 为 migration_pending 占位,assetVersion 未冻结;迁移时冻结版本/hash、来源、适用品类与消费它的 prompt 版本;完成前不得新增 live 消费。",
|
||||
"gac-shanhai-xingji": "active:创始人于 2026-07-27 正式签认地图1《裂谷原》完整20分钟纵切版,音频批准范围为 SFX-only。artifactHash 是同日从当前源重新构建 dist/shanhai-bundle.js 后计算的真实 SHA-256;地图2–5、29兽正式素材与6首BGM属于后续完整版,不在本记录承诺范围。consumerRef 对应 cheap_studio.py → cheap_verify.build_v3_reference_asset_generation_constraints 的生成运行时参照资产注入门。",
|
||||
"gac-shanhai-xunyi-lu": "candidate:designIntent 尚未追认(designRef 为 null);assetRef/artifactHash 为 candidate 占位(资产树待定位绑定);签认前没有正式 game_content_gold 记录。",
|
||||
"gac-yeshi-yitiaojie": "candidate:designIntent 尚未追认(designRef 为 null);assetRef/artifactHash 为 candidate 占位(资产树待定位绑定);签认前没有正式 game_content_gold 记录。"
|
||||
}
|
||||
}
|
||||
10
contracts/play-loop/reference-asset-release.initial.json
Normal file
10
contracts/play-loop/reference-asset-release.initial.json
Normal file
@ -0,0 +1,10 @@
|
||||
{
|
||||
"schemaVersion": "ReferenceAssetRelease/1",
|
||||
"releaseId": "survivor-gold-release-2026-07-27",
|
||||
"registryRef": "contracts/play-loop/reference-asset-registry.v2.initial.json",
|
||||
"registryHash": "57c82adcef617fdb929d056733a395f7c4133636eb16bd38926637857b331b78",
|
||||
"policyRef": "contracts/play-loop/reference-asset-consumption-policy.initial.json",
|
||||
"policyHash": "695b6143fec3c7b30d7be5f5ce088ccb1c1add97899dae9f167bd3a9524d8bbf",
|
||||
"verifierVersion": "reference-asset-verifier/1.0.0",
|
||||
"trustedRootId": "wanxiang-reference-assets-root-v1"
|
||||
}
|
||||
43
contracts/play-loop/reference-asset-release.schema.json
Normal file
43
contracts/play-loop/reference-asset-release.schema.json
Normal file
@ -0,0 +1,43 @@
|
||||
{
|
||||
"$schema": "https://json-schema.org/draft/2020-12/schema",
|
||||
"$id": "https://wanxiang.ai/contracts/play-loop/reference-asset-release.schema.json",
|
||||
"title": "ReferenceAssetReleaseV1",
|
||||
"description": "ReferenceAssetRelease/1。将 ReferenceAssetRegistry/2、ReferenceAssetConsumptionPolicy/1、验证器版本和可信根绑定为不可变 release 身份。",
|
||||
"type": "object",
|
||||
"required": [
|
||||
"schemaVersion", "releaseId", "registryRef", "registryHash",
|
||||
"policyRef", "policyHash", "verifierVersion", "trustedRootId"
|
||||
],
|
||||
"properties": {
|
||||
"schemaVersion": { "const": "ReferenceAssetRelease/1" },
|
||||
"releaseId": {
|
||||
"type": "string",
|
||||
"pattern": "^[A-Za-z0-9_][A-Za-z0-9._-]*$"
|
||||
},
|
||||
"registryRef": { "$ref": "#/$defs/relativePath" },
|
||||
"registryHash": { "$ref": "#/$defs/sha256" },
|
||||
"policyRef": { "$ref": "#/$defs/relativePath" },
|
||||
"policyHash": { "$ref": "#/$defs/sha256" },
|
||||
"verifierVersion": {
|
||||
"type": "string",
|
||||
"pattern": "^[A-Za-z][A-Za-z0-9._-]*/[0-9]+\\.[0-9]+\\.[0-9]+$"
|
||||
},
|
||||
"trustedRootId": {
|
||||
"type": "string",
|
||||
"pattern": "^[A-Za-z0-9_][A-Za-z0-9._:-]*$"
|
||||
}
|
||||
},
|
||||
"additionalProperties": false,
|
||||
"$defs": {
|
||||
"sha256": {
|
||||
"type": "string",
|
||||
"pattern": "^[0-9a-f]{64}$"
|
||||
},
|
||||
"relativePath": {
|
||||
"type": "string",
|
||||
"minLength": 1,
|
||||
"maxLength": 1024,
|
||||
"pattern": "^(?!/)(?!.*//)(?!.*(?:^|/)\\.{1,2}(?:/|$))(?!.*\\\\).+$"
|
||||
}
|
||||
}
|
||||
}
|
||||
@ -0,0 +1,68 @@
|
||||
{
|
||||
"$schema": "https://json-schema.org/draft/2020-12/schema",
|
||||
"$id": "https://wanxiang.ai/contracts/play-loop/reference-asset-verification-receipt.schema.json",
|
||||
"title": "ReferenceAssetVerificationReceiptV1",
|
||||
"description": "ReferenceAssetVerificationReceipt/1。冻结可信 registry、policy、verifier、root、record 和双身份制品/清单的完整验证闭包;expected 与 observed 的 registry/artifact/manifest hash 必须由语义层逐项相等。",
|
||||
"type": "object",
|
||||
"required": [
|
||||
"schemaVersion", "receiptId", "releaseRef", "registryVersion",
|
||||
"expectedRegistryHash", "observedRegistryHash", "policyHash", "verifierVersion",
|
||||
"trustedRootId", "recordId", "role", "consumerRef", "artifactRef",
|
||||
"consumptionManifestRef", "expected", "observed", "finalSnapshotHash"
|
||||
],
|
||||
"properties": {
|
||||
"schemaVersion": { "const": "ReferenceAssetVerificationReceipt/1" },
|
||||
"receiptId": {
|
||||
"type": "string",
|
||||
"pattern": "^[A-Za-z0-9_][A-Za-z0-9._-]*$"
|
||||
},
|
||||
"releaseRef": { "$ref": "#/$defs/relativePath" },
|
||||
"registryVersion": { "type": "string", "minLength": 1 },
|
||||
"expectedRegistryHash": { "$ref": "#/$defs/sha256" },
|
||||
"observedRegistryHash": { "$ref": "#/$defs/sha256" },
|
||||
"policyHash": { "$ref": "#/$defs/sha256" },
|
||||
"verifierVersion": {
|
||||
"type": "string",
|
||||
"pattern": "^[A-Za-z][A-Za-z0-9._-]*/[0-9]+\\.[0-9]+\\.[0-9]+$"
|
||||
},
|
||||
"trustedRootId": {
|
||||
"type": "string",
|
||||
"pattern": "^[A-Za-z0-9_][A-Za-z0-9._:-]*$"
|
||||
},
|
||||
"recordId": {
|
||||
"type": "string",
|
||||
"pattern": "^[A-Za-z0-9_][A-Za-z0-9._-]*$"
|
||||
},
|
||||
"role": {
|
||||
"enum": ["harness_fixture", "prompt_eval_gold", "generation_exemplar", "game_content_gold"]
|
||||
},
|
||||
"consumerRef": { "type": "string", "minLength": 1 },
|
||||
"artifactRef": { "$ref": "#/$defs/relativePath" },
|
||||
"consumptionManifestRef": { "$ref": "#/$defs/relativePath" },
|
||||
"expected": { "$ref": "#/$defs/hashColumns" },
|
||||
"observed": { "$ref": "#/$defs/hashColumns" },
|
||||
"finalSnapshotHash": { "$ref": "#/$defs/sha256" }
|
||||
},
|
||||
"additionalProperties": false,
|
||||
"$defs": {
|
||||
"hashColumns": {
|
||||
"type": "object",
|
||||
"required": ["artifactHash", "consumptionManifestHash"],
|
||||
"properties": {
|
||||
"artifactHash": { "$ref": "#/$defs/sha256" },
|
||||
"consumptionManifestHash": { "$ref": "#/$defs/sha256" }
|
||||
},
|
||||
"additionalProperties": false
|
||||
},
|
||||
"sha256": {
|
||||
"type": "string",
|
||||
"pattern": "^[0-9a-f]{64}$"
|
||||
},
|
||||
"relativePath": {
|
||||
"type": "string",
|
||||
"minLength": 1,
|
||||
"maxLength": 1024,
|
||||
"pattern": "^(?!/)(?!.*//)(?!.*(?:^|/)\\.{1,2}(?:/|$))(?!.*\\\\).+$"
|
||||
}
|
||||
}
|
||||
}
|
||||
@ -0,0 +1,5 @@
|
||||
{
|
||||
"_base": "../valid/02-with-consumed-assets.json",
|
||||
"_why": "消费溯源三元组 recordId+role+artifactHash 缺一不可;缺 hash 即无法对账漂移",
|
||||
"_delete": ["/consumedReferenceAssets/0/artifactHash"]
|
||||
}
|
||||
@ -0,0 +1,7 @@
|
||||
{
|
||||
"_base": "../valid/02-with-consumed-assets.json",
|
||||
"_why": "role 快照枚举必须与 ReferenceAssetRecord/1.role 四值一致",
|
||||
"_set": {
|
||||
"/consumedReferenceAssets/0/role": "game_gold"
|
||||
}
|
||||
}
|
||||
@ -0,0 +1,16 @@
|
||||
{
|
||||
"schemaVersion": "acceptance-provenance/3",
|
||||
"gameId": "sample-puzzle-v3-base-001",
|
||||
"briefHash": "8728bf7475809ad48aaedb2a26601f030d94981e186f4e2818628d9962f59696",
|
||||
"genre": "puzzle",
|
||||
"templateRoute": "_template-puzzle",
|
||||
"proofProfileId": "puzzle.match-board",
|
||||
"proofRegistryVersion": "2026-07-15.v3",
|
||||
"taskBindingHash": "aaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaa",
|
||||
"interactionBinding": null,
|
||||
"sourceArtifactHash": null,
|
||||
"parentAcceptanceRequestHash": null,
|
||||
"repairOrdinal": 0,
|
||||
"acceptanceRequestHash": "fc1769d63f64dba013af63268c8b6053d2205a5513422e95d9456bff802d564c",
|
||||
"artifactHash": "3333333333333333333333333333333333333333333333333333333333333333"
|
||||
}
|
||||
@ -0,0 +1,31 @@
|
||||
{
|
||||
"schemaVersion": "acceptance-provenance/3",
|
||||
"gameId": "sample-puzzle-v3-refs-001",
|
||||
"briefHash": "8728bf7475809ad48aaedb2a26601f030d94981e186f4e2818628d9962f59696",
|
||||
"genre": "puzzle",
|
||||
"templateRoute": "_template-puzzle",
|
||||
"proofProfileId": "puzzle.match-board",
|
||||
"proofRegistryVersion": "2026-07-15.v3",
|
||||
"taskBindingHash": "aaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaa",
|
||||
"interactionBinding": null,
|
||||
"sourceArtifactHash": null,
|
||||
"parentAcceptanceRequestHash": null,
|
||||
"repairOrdinal": 0,
|
||||
"acceptanceRequestHash": "fc1769d63f64dba013af63268c8b6053d2205a5513422e95d9456bff802d564c",
|
||||
"artifactHash": "3333333333333333333333333333333333333333333333333333333333333333",
|
||||
"designRef": "gac-shanhai-xingji",
|
||||
"referenceAssetRecordIds": ["_template-puzzle", "gold-m3-gem-r3"],
|
||||
"consumerRef": "cheap-worker.run_acceptance_v3@2026-07-25",
|
||||
"consumedReferenceAssets": [
|
||||
{
|
||||
"recordId": "_template-puzzle",
|
||||
"role": "generation_exemplar",
|
||||
"artifactHash": "9ceee42869560d5806184544ff9668bc8c56a4b582a74dc6bb8fab9511281fa9"
|
||||
},
|
||||
{
|
||||
"recordId": "gold-m3-gem-r3",
|
||||
"role": "harness_fixture",
|
||||
"artifactHash": "94075fb645952bd068c8429042a0e82247ffe66e72c48951e1af68603d71f4fc"
|
||||
}
|
||||
]
|
||||
}
|
||||
@ -0,0 +1,4 @@
|
||||
{
|
||||
"_base": "../valid/01-with-verification-receipt.json",
|
||||
"_delete": ["/referenceAssetVerificationReceipts"]
|
||||
}
|
||||
@ -0,0 +1,4 @@
|
||||
{
|
||||
"_base": "../valid/01-with-verification-receipt.json",
|
||||
"_set": {"/referenceAssetVerificationReceipts": []}
|
||||
}
|
||||
@ -0,0 +1,4 @@
|
||||
{
|
||||
"_base": "../valid/01-with-verification-receipt.json",
|
||||
"_set": {"/referenceAssetVerificationReceipts/0/hash": "wrong"}
|
||||
}
|
||||
@ -0,0 +1,4 @@
|
||||
{
|
||||
"_base": "../valid/01-with-verification-receipt.json",
|
||||
"_set": {"/referenceAssetVerificationReceipts/0/ref": "../receipt.json"}
|
||||
}
|
||||
@ -0,0 +1,4 @@
|
||||
{
|
||||
"_base": "../valid/02-repair-with-receipt.json",
|
||||
"_set": {"/referenceAssetVerificationReceipts/0/verified": true}
|
||||
}
|
||||
@ -0,0 +1,4 @@
|
||||
{
|
||||
"_base": "../valid/01-with-verification-receipt.json",
|
||||
"_set": {"/genre": "narrative"}
|
||||
}
|
||||
@ -0,0 +1,4 @@
|
||||
{
|
||||
"_base": "../valid/01-with-verification-receipt.json",
|
||||
"_set": {"/templateRoute": "_template-story"}
|
||||
}
|
||||
@ -0,0 +1,4 @@
|
||||
{
|
||||
"_base": "../valid/01-with-verification-receipt.json",
|
||||
"_set": {"/proofRegistryVersion": "bogus"}
|
||||
}
|
||||
@ -0,0 +1,6 @@
|
||||
{
|
||||
"_base": "../valid/03-verified-match3-with-binding.json",
|
||||
"_set": {
|
||||
"/interactionBinding/taskBindingHash": "bbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbb"
|
||||
}
|
||||
}
|
||||
@ -0,0 +1,32 @@
|
||||
{
|
||||
"schemaVersion": "acceptance-provenance/4",
|
||||
"gameId": "shanhai-xingji",
|
||||
"briefHash": "aaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaa",
|
||||
"genre": "heritage",
|
||||
"templateRoute": "_template-feiyi",
|
||||
"proofProfileId": "heritage.ordered-craft",
|
||||
"proofRegistryVersion": "2026-07-15.v3",
|
||||
"taskBindingHash": "bbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbb",
|
||||
"interactionBinding": null,
|
||||
"sourceArtifactHash": null,
|
||||
"parentAcceptanceRequestHash": null,
|
||||
"repairOrdinal": 0,
|
||||
"acceptanceRequestHash": "cccccccccccccccccccccccccccccccccccccccccccccccccccccccccccccccc",
|
||||
"artifactHash": "dddddddddddddddddddddddddddddddddddddddddddddddddddddddddddddddd",
|
||||
"designRef": "docs/agent-specs/shanhai-design.md",
|
||||
"referenceAssetRecordIds": ["gac-shanhai-xingji"],
|
||||
"consumerRef": "generation-runtime@reference-assets/2",
|
||||
"consumedReferenceAssets": [
|
||||
{
|
||||
"recordId": "gac-shanhai-xingji",
|
||||
"role": "game_content_gold",
|
||||
"artifactHash": "1c760811ec435fe0f3b5ba79aa8c4fcc44119e019b3e1240ce2c51e55e25870b"
|
||||
}
|
||||
],
|
||||
"referenceAssetVerificationReceipts": [
|
||||
{
|
||||
"ref": "contracts/play-loop/samples/reference-asset-verification-receipt/valid/01-matching-hashes.json",
|
||||
"hash": "9e127640db31aeae4ecb2262e35632a31efb773ef0c3c204ee0ffc52d9e5ee8c"
|
||||
}
|
||||
]
|
||||
}
|
||||
@ -0,0 +1,31 @@
|
||||
{
|
||||
"schemaVersion": "acceptance-provenance/4",
|
||||
"gameId": "fixture-puzzle",
|
||||
"briefHash": "1111111111111111111111111111111111111111111111111111111111111111",
|
||||
"genre": "puzzle",
|
||||
"templateRoute": "_template-puzzle",
|
||||
"proofProfileId": "puzzle.match-board",
|
||||
"proofRegistryVersion": "2026-07-15.v3",
|
||||
"taskBindingHash": "2222222222222222222222222222222222222222222222222222222222222222",
|
||||
"interactionBinding": null,
|
||||
"sourceArtifactHash": "3333333333333333333333333333333333333333333333333333333333333333",
|
||||
"parentAcceptanceRequestHash": "4444444444444444444444444444444444444444444444444444444444444444",
|
||||
"repairOrdinal": 1,
|
||||
"acceptanceRequestHash": "5555555555555555555555555555555555555555555555555555555555555555",
|
||||
"artifactHash": "6666666666666666666666666666666666666666666666666666666666666666",
|
||||
"referenceAssetRecordIds": ["gac-shanhai-xingji"],
|
||||
"consumerRef": "generation-runtime@reference-assets/2",
|
||||
"consumedReferenceAssets": [
|
||||
{
|
||||
"recordId": "gac-shanhai-xingji",
|
||||
"role": "game_content_gold",
|
||||
"artifactHash": "1c760811ec435fe0f3b5ba79aa8c4fcc44119e019b3e1240ce2c51e55e25870b"
|
||||
}
|
||||
],
|
||||
"referenceAssetVerificationReceipts": [
|
||||
{
|
||||
"ref": "contracts/play-loop/samples/reference-asset-verification-receipt/valid/01-matching-hashes.json",
|
||||
"hash": "9e127640db31aeae4ecb2262e35632a31efb773ef0c3c204ee0ffc52d9e5ee8c"
|
||||
}
|
||||
]
|
||||
}
|
||||
@ -0,0 +1,17 @@
|
||||
{
|
||||
"_base": "01-with-verification-receipt.json",
|
||||
"_set": {
|
||||
"/gameId": "sample-match3-v4-verified-001",
|
||||
"/genre": "puzzle",
|
||||
"/templateRoute": "_template-puzzle",
|
||||
"/proofProfileId": "puzzle.match-board",
|
||||
"/taskBindingHash": "aaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaa",
|
||||
"/interactionBinding": {
|
||||
"schemaVersion": "InteractionBinding/1",
|
||||
"interactionProfileId": "match3.orthogonal-swap-v1",
|
||||
"interactionRegistryVersion": "2026-07-15.v1",
|
||||
"taskBindingHash": "aaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaa",
|
||||
"interactionBindingHash": "1c9efaf79571d5e2d3c9da033cf4d79678b4735ddb4c9ab29d8384d17ebcd3ce"
|
||||
}
|
||||
}
|
||||
}
|
||||
@ -0,0 +1,7 @@
|
||||
{
|
||||
"_base": "../valid/01-base-without-reference-refs.json",
|
||||
"_why": "v3 只开口三个声明字段;未登记字段仍被 additionalProperties=false 拒绝",
|
||||
"_set": {
|
||||
"/referenceAssetBundle": "anything"
|
||||
}
|
||||
}
|
||||
@ -0,0 +1,7 @@
|
||||
{
|
||||
"_base": "../valid/02-with-reference-refs.json",
|
||||
"_why": "referenceAssetRecordIds 元素必须形如 ReferenceAssetRecord/1.recordId;含空格/中文即拦",
|
||||
"_set": {
|
||||
"/referenceAssetRecordIds/1": "gold m3-宝石"
|
||||
}
|
||||
}
|
||||
@ -0,0 +1,8 @@
|
||||
{
|
||||
"_base": "../valid/02-with-reference-refs.json",
|
||||
"_why": "v3 不放松 v2 的修回约束:repairOrdinal=1 必须同时绑定父请求与原产物",
|
||||
"_set": {
|
||||
"/repairOrdinal": 1,
|
||||
"/sourceArtifactHash": "3333333333333333333333333333333333333333333333333333333333333333"
|
||||
}
|
||||
}
|
||||
@ -0,0 +1,14 @@
|
||||
{
|
||||
"schemaVersion": "acceptance-request/3",
|
||||
"gameId": "sample-puzzle-v3-base-001",
|
||||
"briefHash": "8728bf7475809ad48aaedb2a26601f030d94981e186f4e2818628d9962f59696",
|
||||
"genre": "puzzle",
|
||||
"templateRoute": "_template-puzzle",
|
||||
"proofProfileId": "puzzle.match-board",
|
||||
"proofRegistryVersion": "2026-07-15.v3",
|
||||
"taskBindingHash": "aaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaa",
|
||||
"interactionBinding": null,
|
||||
"sourceArtifactHash": null,
|
||||
"parentAcceptanceRequestHash": null,
|
||||
"repairOrdinal": 0
|
||||
}
|
||||
@ -0,0 +1,17 @@
|
||||
{
|
||||
"schemaVersion": "acceptance-request/3",
|
||||
"gameId": "sample-puzzle-v3-refs-001",
|
||||
"briefHash": "8728bf7475809ad48aaedb2a26601f030d94981e186f4e2818628d9962f59696",
|
||||
"genre": "puzzle",
|
||||
"templateRoute": "_template-puzzle",
|
||||
"proofProfileId": "puzzle.match-board",
|
||||
"proofRegistryVersion": "2026-07-15.v3",
|
||||
"taskBindingHash": "aaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaa",
|
||||
"interactionBinding": null,
|
||||
"sourceArtifactHash": null,
|
||||
"parentAcceptanceRequestHash": null,
|
||||
"repairOrdinal": 0,
|
||||
"designRef": "gac-shanhai-xingji",
|
||||
"referenceAssetRecordIds": ["_template-puzzle", "gold-m3-gem-r3"],
|
||||
"consumerRef": "cheap-worker.run_acceptance_v3@2026-07-25"
|
||||
}
|
||||
@ -0,0 +1,8 @@
|
||||
{
|
||||
"_base": "02-with-reference-refs.json",
|
||||
"_set": {
|
||||
"/sourceArtifactHash": "3333333333333333333333333333333333333333333333333333333333333333",
|
||||
"/parentAcceptanceRequestHash": "fc1769d63f64dba013af63268c8b6053d2205a5513422e95d9456bff802d564c",
|
||||
"/repairOrdinal": 1
|
||||
}
|
||||
}
|
||||
@ -0,0 +1,4 @@
|
||||
{
|
||||
"_base": "../valid/01-canonical-manifest.json",
|
||||
"_set": {"/entries/1/path": "dist/index.html"}
|
||||
}
|
||||
@ -0,0 +1,4 @@
|
||||
{
|
||||
"_base": "../valid/01-canonical-manifest.json",
|
||||
"_set": {"/entries/0/path": "dist/z.js"}
|
||||
}
|
||||
@ -0,0 +1,4 @@
|
||||
{
|
||||
"_base": "../valid/02-single-entry.json",
|
||||
"_set": {"/entries/0/path": "assets/e\u0301.js"}
|
||||
}
|
||||
@ -0,0 +1,4 @@
|
||||
{
|
||||
"_base": "../valid/02-single-entry.json",
|
||||
"_set": {"/entries/0/path": "../outside.js"}
|
||||
}
|
||||
@ -0,0 +1,4 @@
|
||||
{
|
||||
"_base": "../valid/02-single-entry.json",
|
||||
"_set": {"/entries/0/sha256": "not-a-sha256"}
|
||||
}
|
||||
@ -0,0 +1,4 @@
|
||||
{
|
||||
"_base": "../valid/02-single-entry.json",
|
||||
"_set": {"/entries/0/size": -1}
|
||||
}
|
||||
@ -0,0 +1,4 @@
|
||||
{
|
||||
"_base": "../valid/02-single-entry.json",
|
||||
"_set": {"/entries/0/extra": true}
|
||||
}
|
||||
@ -0,0 +1,4 @@
|
||||
{
|
||||
"_base": "../valid/02-single-entry.json",
|
||||
"_set": {"/canonicalization": "pretty-json"}
|
||||
}
|
||||
@ -0,0 +1 @@
|
||||
{"canonicalization":"reference-asset-consumption-manifest/1","entries":[{"path":"bundle.js","sha256":"3333333333333333333333333333333333333333333333333333333333333333","size":0}],"manifestId":"fixture-manifest-v1","schemaVersion":"ReferenceAssetConsumptionManifest/1"}
|
||||
@ -0,0 +1 @@
|
||||
{"canonicalization":"reference-asset-consumption-manifest/1","entries":[{"path":"bundle.js","sha256":"3333333333333333333333333333333333333333333333333333333333333333","size":0}],"manifestId":"fixture-manifest-v1","schemaVersion":"ReferenceAssetConsumptionManifest/1"}
|
||||
@ -0,0 +1 @@
|
||||
{"schemaVersion":"ReferenceAssetConsumptionManifest/1","manifestId":"fixture-manifest-v1","canonicalization":"reference-asset-consumption-manifest/1","entries":[{"path":"bundle.js","size":0,"sha256":"3333333333333333333333333333333333333333333333333333333333333333"}]}
|
||||
@ -0,0 +1 @@
|
||||
{"canonicalization":"reference-asset-consumption-manifest/1","entries":[{"path":"assets/\u96ea.png","sha256":"4444444444444444444444444444444444444444444444444444444444444444","size":1}],"manifestId":"fixture-manifest-unicode-v1","schemaVersion":"ReferenceAssetConsumptionManifest/1"}
|
||||
@ -0,0 +1 @@
|
||||
{"canonicalization":"reference-asset-consumption-manifest/1","entries":[{"path":"dist/index.html","sha256":"1111111111111111111111111111111111111111111111111111111111111111","size":1024},{"path":"dist/main.js","sha256":"2222222222222222222222222222222222222222222222222222222222222222","size":2048}],"manifestId":"survivor-gold-v1-manifest","schemaVersion":"ReferenceAssetConsumptionManifest/1"}
|
||||
@ -0,0 +1 @@
|
||||
{"canonicalization":"reference-asset-consumption-manifest/1","entries":[{"path":"bundle.js","sha256":"3333333333333333333333333333333333333333333333333333333333333333","size":0}],"manifestId":"fixture-manifest-v1","schemaVersion":"ReferenceAssetConsumptionManifest/1"}
|
||||
@ -0,0 +1,4 @@
|
||||
{
|
||||
"_base": "../valid/01-survivor-gold-v1.json",
|
||||
"_set": {"/autoSelect": true}
|
||||
}
|
||||
@ -0,0 +1,4 @@
|
||||
{
|
||||
"_base": "../valid/01-survivor-gold-v1.json",
|
||||
"_set": {"/mode": "auto"}
|
||||
}
|
||||
@ -0,0 +1,4 @@
|
||||
{
|
||||
"_base": "../valid/01-survivor-gold-v1.json",
|
||||
"_set": {"/recordId": "another-gold"}
|
||||
}
|
||||
@ -0,0 +1,4 @@
|
||||
{
|
||||
"_base": "../valid/01-survivor-gold-v1.json",
|
||||
"_set": {"/fallback": "latest"}
|
||||
}
|
||||
@ -0,0 +1,10 @@
|
||||
{
|
||||
"schemaVersion": "ReferenceAssetConsumptionPolicy/1",
|
||||
"policyId": "survivor-gold-v1",
|
||||
"recordId": "gac-shanhai-xingji",
|
||||
"role": "game_content_gold",
|
||||
"consumerRef": "generation-runtime@reference-assets/2",
|
||||
"route": "survivor-gold",
|
||||
"autoSelect": false,
|
||||
"mode": "frozen_preflight"
|
||||
}
|
||||
@ -0,0 +1,4 @@
|
||||
{
|
||||
"_base": "../valid/01-active-trusted-gold.json",
|
||||
"_delete": ["/artifactRef"]
|
||||
}
|
||||
@ -0,0 +1,4 @@
|
||||
{
|
||||
"_base": "../valid/01-active-trusted-gold.json",
|
||||
"_set": {"/consumptionManifestHash": null}
|
||||
}
|
||||
@ -0,0 +1,4 @@
|
||||
{
|
||||
"_base": "../valid/03-retired-with-history.json",
|
||||
"_delete": ["/signedAt"]
|
||||
}
|
||||
@ -0,0 +1,4 @@
|
||||
{
|
||||
"_base": "../valid/02-candidate-unbound.json",
|
||||
"_set": {"/unexpectedField": true}
|
||||
}
|
||||
@ -0,0 +1,4 @@
|
||||
{
|
||||
"_base": "../valid/01-active-trusted-gold.json",
|
||||
"_set": {"/consumptionManifestHash": "not-a-sha256"}
|
||||
}
|
||||
@ -0,0 +1,4 @@
|
||||
{
|
||||
"_base": "../valid/01-active-trusted-gold.json",
|
||||
"_set": {"/artifactRef": "assets/e\u0301.js"}
|
||||
}
|
||||
@ -0,0 +1,4 @@
|
||||
{
|
||||
"_base": "../valid/01-active-trusted-gold.json",
|
||||
"_set": {"/artifactRef": "../outside.js"}
|
||||
}
|
||||
@ -0,0 +1,4 @@
|
||||
{
|
||||
"_base": "../valid/02-candidate-unbound.json",
|
||||
"_set": {"/schemaVersion": "ReferenceAssetRecord/1"}
|
||||
}
|
||||
@ -0,0 +1,22 @@
|
||||
{
|
||||
"schemaVersion": "ReferenceAssetRecord/2",
|
||||
"recordId": "gac-shanhai-xingji",
|
||||
"role": "game_content_gold",
|
||||
"lifecycleStatus": "active",
|
||||
"assetRef": "game-runtime/games/shanhai-xingji",
|
||||
"assetVersion": "map1-vertical-slice-r1",
|
||||
"artifactHash": "1c760811ec435fe0f3b5ba79aa8c4fcc44119e019b3e1240ce2c51e55e25870b",
|
||||
"consumerRef": "generation-runtime@reference-assets/2",
|
||||
"designRef": [
|
||||
"docs/agent-specs/2026-07-06-北极星顶级线-肉鸽割草-开发设计书.md",
|
||||
"docs/agent-specs/2026-07-06-山海宇宙设定与美术音频管线-选型材料.md"
|
||||
],
|
||||
"evidenceRefs": [
|
||||
"game-runtime/games/shanhai-xingji/evidence/round-17-gold-lock/qa-report.md"
|
||||
],
|
||||
"signedBy": "创始人",
|
||||
"signedAt": "2026-07-27T13:30:50Z",
|
||||
"artifactRef": "game-runtime/games/shanhai-xingji/dist/shanhai-bundle.js",
|
||||
"consumptionManifestRef": "contracts/play-loop/samples/reference-asset-consumption-manifest/valid/01-canonical-manifest.json",
|
||||
"consumptionManifestHash": "2810af6e398f8a7f9c6d970585814fdac1d91f986303f15e27aa3e49928495cd"
|
||||
}
|
||||
@ -0,0 +1,17 @@
|
||||
{
|
||||
"schemaVersion": "ReferenceAssetRecord/2",
|
||||
"recordId": "gac-shanhai-xunyi-lu",
|
||||
"role": "game_content_gold",
|
||||
"lifecycleStatus": "candidate",
|
||||
"assetRef": "pending-binding-fable-shanhai-xunyi-lu",
|
||||
"assetVersion": "unfrozen-2026-07-27",
|
||||
"artifactHash": "17dc264e89ed8541fde97029cb87a7426af69a694cfdc70b87c19da0bf0b6f33",
|
||||
"consumerRef": null,
|
||||
"designRef": null,
|
||||
"evidenceRefs": [],
|
||||
"signedBy": null,
|
||||
"signedAt": null,
|
||||
"artifactRef": null,
|
||||
"consumptionManifestRef": null,
|
||||
"consumptionManifestHash": null
|
||||
}
|
||||
@ -0,0 +1,17 @@
|
||||
{
|
||||
"schemaVersion": "ReferenceAssetRecord/2",
|
||||
"recordId": "gac-retired-gold",
|
||||
"role": "game_content_gold",
|
||||
"lifecycleStatus": "retired",
|
||||
"assetRef": "game-runtime/games/retired-gold",
|
||||
"assetVersion": "r1",
|
||||
"artifactHash": "cccccccccccccccccccccccccccccccccccccccccccccccccccccccccccccccc",
|
||||
"consumerRef": "generation-runtime@reference-assets/1",
|
||||
"designRef": ["docs/agent-specs/retired-gold-design.md"],
|
||||
"evidenceRefs": ["evidence/retired-gold/signoff.md"],
|
||||
"signedBy": "founder",
|
||||
"signedAt": "2026-07-20",
|
||||
"artifactRef": "game-runtime/games/retired-gold/dist/bundle.js",
|
||||
"consumptionManifestRef": "contracts/play-loop/samples/reference-asset-consumption-manifest/valid/01-canonical-manifest.json",
|
||||
"consumptionManifestHash": "dddddddddddddddddddddddddddddddddddddddddddddddddddddddddddddddd"
|
||||
}
|
||||
@ -0,0 +1,17 @@
|
||||
{
|
||||
"schemaVersion": "ReferenceAssetRecord/2",
|
||||
"recordId": "gac-retired-untrusted",
|
||||
"role": "harness_fixture",
|
||||
"lifecycleStatus": "retired",
|
||||
"assetRef": "game-runtime/games/retired-untrusted",
|
||||
"assetVersion": "r0",
|
||||
"artifactHash": "1212121212121212121212121212121212121212121212121212121212121212",
|
||||
"consumerRef": null,
|
||||
"designRef": null,
|
||||
"evidenceRefs": [],
|
||||
"signedBy": null,
|
||||
"signedAt": null,
|
||||
"artifactRef": null,
|
||||
"consumptionManifestRef": null,
|
||||
"consumptionManifestHash": null
|
||||
}
|
||||
@ -0,0 +1,10 @@
|
||||
{
|
||||
"schemaVersion": "ReferenceAssetRecord/2",
|
||||
"recordId": "gac-retired-legacy",
|
||||
"role": "harness_fixture",
|
||||
"lifecycleStatus": "retired",
|
||||
"assetRef": "game-runtime/games/retired-legacy",
|
||||
"assetVersion": "r0",
|
||||
"artifactHash": "1313131313131313131313131313131313131313131313131313131313131313",
|
||||
"evidenceRefs": []
|
||||
}
|
||||
@ -0,0 +1,5 @@
|
||||
{
|
||||
"_base": "../valid/03-active-game-content-gold.json",
|
||||
"_why": "active 必须签认:缺 signedBy 即不满足金标 SoT §7『active 时必填』",
|
||||
"_delete": ["/signedBy"]
|
||||
}
|
||||
@ -0,0 +1,7 @@
|
||||
{
|
||||
"_base": "../valid/01-harness-fixture-migration-pending.json",
|
||||
"_why": "role 只能四值互斥;自造第五类即拦",
|
||||
"_set": {
|
||||
"/role": "game_gold"
|
||||
}
|
||||
}
|
||||
@ -0,0 +1,7 @@
|
||||
{
|
||||
"_base": "../valid/03-active-game-content-gold.json",
|
||||
"_why": "game_content_gold 升 active 必须指向已批准 designIntent(金标 SoT §7 designRef 行)",
|
||||
"_set": {
|
||||
"/designRef": null
|
||||
}
|
||||
}
|
||||
@ -0,0 +1,7 @@
|
||||
{
|
||||
"_base": "../valid/01-harness-fixture-migration-pending.json",
|
||||
"_why": "artifactHash 必须是 64 位小写十六进制;大写/缺位即拦",
|
||||
"_set": {
|
||||
"/artifactHash": "94075FB645952BD0"
|
||||
}
|
||||
}
|
||||
@ -0,0 +1,14 @@
|
||||
{
|
||||
"schemaVersion": "ReferenceAssetRecord/1",
|
||||
"recordId": "gold-m3-gem-r3",
|
||||
"role": "harness_fixture",
|
||||
"lifecycleStatus": "migration_pending",
|
||||
"assetRef": "game-runtime/games/_wg1-gen/gold-m3-gem-r3",
|
||||
"assetVersion": "r3",
|
||||
"artifactHash": "94075fb645952bd068c8429042a0e82247ffe66e72c48951e1af68603d71f4fc",
|
||||
"consumerRef": null,
|
||||
"designRef": null,
|
||||
"evidenceRefs": [],
|
||||
"signedBy": null,
|
||||
"signedAt": null
|
||||
}
|
||||
@ -0,0 +1,17 @@
|
||||
{
|
||||
"schemaVersion": "ReferenceAssetRecord/1",
|
||||
"recordId": "gac-shanhai-xingji",
|
||||
"role": "game_content_gold",
|
||||
"lifecycleStatus": "candidate",
|
||||
"assetRef": "game-runtime/games/shanhai-xingji",
|
||||
"assetVersion": "unfrozen-2026-07-25",
|
||||
"artifactHash": "eab2cab4b8ef9e8431a30194104d1750059c3b804cb74e91bc57e72315c7c8f8",
|
||||
"consumerRef": null,
|
||||
"designRef": [
|
||||
"docs/agent-specs/2026-07-06-北极星顶级线-肉鸽割草-开发设计书.md",
|
||||
"docs/agent-specs/2026-07-06-山海宇宙设定与美术音频管线-选型材料.md"
|
||||
],
|
||||
"evidenceRefs": [],
|
||||
"signedBy": null,
|
||||
"signedAt": null
|
||||
}
|
||||
@ -0,0 +1,14 @@
|
||||
{
|
||||
"schemaVersion": "ReferenceAssetRecord/1",
|
||||
"recordId": "gac-example-active-gold",
|
||||
"role": "game_content_gold",
|
||||
"lifecycleStatus": "active",
|
||||
"assetRef": "game-runtime/games/example-gold",
|
||||
"assetVersion": "m6-frozen",
|
||||
"artifactHash": "eab2cab4b8ef9e8431a30194104d1750059c3b804cb74e91bc57e72315c7c8f8",
|
||||
"consumerRef": "survivor.brief-compiler@1.0.0",
|
||||
"designRef": ["docs/agent-specs/2026-07-06-北极星顶级线-肉鸽割草-开发设计书.md"],
|
||||
"evidenceRefs": ["evidence/example-gold/nine-dimension-signoff.md"],
|
||||
"signedBy": "founder",
|
||||
"signedAt": "2026-08-01"
|
||||
}
|
||||
@ -0,0 +1,4 @@
|
||||
{
|
||||
"_base": "../valid/01-trusted-release-registry.json",
|
||||
"_set": {"/records/1/recordId": "gac-shanhai-xingji"}
|
||||
}
|
||||
@ -0,0 +1,4 @@
|
||||
{
|
||||
"_base": "../valid/01-trusted-release-registry.json",
|
||||
"_set": {"/registryVersion": "2026-07-27.migration-list"}
|
||||
}
|
||||
Some files were not shown because too many files have changed in this diff Show More
Loading…
x
Reference in New Issue
Block a user