每个失败子检查透传 harness 实测 detail + 一句为什么(_RICH_CHECK_WHY,对齐便宜档 _CHEAP_GATE_HINTS 范式);经济门失败暴露三个数(实测 coins 终值/游戏内赢线 coinsTarget/ 验收外生 winThreshold),游戏内赢线低于验收地板直接指出结构性错配;静态 economyConsistent 差额折进反馈(静态可达而真玩挂→指向耐心/时序/节奏;静态死局→先修数值);H_progress 失败 喂回本局 play-spec 真玩断言清单(path/op/why)。run_gates 返回加性携带 gameId 供反馈组装 读盘;判定逻辑(passed/failed_gates/decision)与 harness 零变化;便宜档不动。 单测 9 新增 + 全套 70 绿。 Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
540 lines
34 KiB
Python
540 lines
34 KiB
Python
"""service/control_plane.py —— tier2 富游戏自治线 · 服务态【有界 resume 控制面】(本批皇冠;B1)。
|
|
|
|
【它补的是哪个 followup】
|
|
CLI 主链 worker.agent_loop.studio.run_studio(:404-458)在 Agent 的 ReAct 循环【之外】套了一圈
|
|
有界 resume:AgentScope 2.0.2 原生 ReAct 在「模型产出无 tool_call 的纯文本回合」即退出
|
|
(agentscope/agent/_agent.py:612)——实测 M3 调一次 run_gates 看到 decision=fix 就产空文本收尾、
|
|
循环退出(看一次 verdict 就放弃、不自纠)。run_studio 用「for attempt in range(max_resumes+1):
|
|
reply→checkpoint→门绿则踹 finish / 否则带 verdict 反馈 resume 续修」兜住了它。
|
|
但**服务态(service/app.py 的 create_app :8200)没有这道循环**:它的 chat 是每回合一次、
|
|
fire-and-forget(_service/_chat.py),由消费方(控制面)决定要不要再 POST /chat。service/app.py
|
|
文件头(:34-42)与 service/bootstrap.py(:25-32)都把它列为 followup。本模块就是来补这个 followup——
|
|
在服务消费方实现等价 CLI max_resumes 的有界 resume 控制面,**绝不重写任何 REST 调用 / 生成逻辑**,
|
|
只 import 复用 bootstrap(REST 原语)+ run(纯代码门 / 落库)。
|
|
|
|
【与 CLI 态的关键差异 —— 为什么不靠内存 session 判收敛】
|
|
服务态九工具(scaffold_init/write_source/build/run_gates/finish)由 _tier2_tools_factory 在【每个
|
|
chat 回合】工厂新建一个 Tier2Session(app.py:_tier2_tools_factory:129-154,把 session_id 当 game_id),
|
|
其内存态 session.finished / last_verdict **跨回合不保留**(每回合一个新 Tier2Session 实例)。
|
|
所以本控制面【不靠内存 session】判收敛,而是【每回合 chat run 结束后独立跑 run.run_gates 判门】——
|
|
门是机器判的(run_gates 纯代码、零 LLM、判 on-disk 源工程),这正是项目「门机器判、控制面绝不
|
|
自评翻绿」的哲学。落库同理:据 on-disk 源工程重建 source_project + run.persist_source_project,
|
|
不依赖内存 session.finished。
|
|
|
|
【game_id ↔ run_gates 评门目标的绑定(诚实写清)】
|
|
服务态九工具实际写文件的工程目录 = run._workdir(<game_id>),而九工具的 game_id 绑的是 AgentScope
|
|
的 session_id(app.py:_tier2_tools_factory:150「game_id 绑 session_id」)。故本控制面评门时,
|
|
**run_gates 的 game_id 必须用 start_new_game 返回的 session_id**——这样评的就是九工具刚写的那个工程目录,
|
|
二者指向同一物理目录(game-runtime/games/_tier2-gen/<session_id>),不会评错对象。
|
|
|
|
【收敛 / 停机判据(对齐 run_studio:436-449)】
|
|
外层 for attempt in range(max_resumes+1):
|
|
a. 等本回合 chat run 结束:消费 SSE GET /sessions/{session_id}/stream,读到本轮 reply 的
|
|
REPLY_END 事件(agentscope/event/_event.py:EventType.REPLY_END;也认 EXCEED_MAX_ITERS 为结束)
|
|
即视为「这轮 agent 跑完了」。带总超时兜底;SSE 断流也据当前 on-disk 状态继续评门,绝不卡死。
|
|
b. 独立评门:res = run.run_gates(session_id, play_spec);v = res['verdict']。
|
|
c. 门绿(v.decision=='accept' 且 layerResults.L1.passed)→ 据 on-disk 源工程落库,break 成功。
|
|
d. attempt>=max_resumes → 预算耗尽,break(停机原因 budget_exhausted)。
|
|
e. 否则:fb = run.verdict_feedback(v);bootstrap.resume_session(..., feedback=续跑指令),回到 a。
|
|
熔断/异常(resume 抛 / SSE 长时间无事件 / run_gates 异常)→ best-effort 记日志、落部分结果返回,绝不静默吞。
|
|
|
|
【惰性 import 红线】本模块顶层【只 import 标准库】。bootstrap/run/store/genconfig 这些包内模块顶层会
|
|
牵出 agentscope / httpx(6c6g 未装),故对它们一律在函数体内 import,保证 6c6g 能
|
|
`python -m py_compile service/control_plane.py` 且能裸 `import service.control_plane` 不炸。
|
|
真跑门(run_gates / chrome / esbuild)只在 mini-desktop;6c6g 上 run_gates 返回结构化失败(诚实失败)。
|
|
"""
|
|
|
|
from __future__ import annotations
|
|
|
|
import json
|
|
import time
|
|
from typing import Any
|
|
|
|
|
|
# ── SSE 事件常量(源码核验 agentscope/event/_event.py:EventType)──
|
|
# EventBase.model_config use_enum_values=True ⇒ model_dump(mode="json") 的 "type" 字段是字符串值。
|
|
# ReplyEndEvent.type = EventType.REPLY_END → 序列化后 type=="REPLY_END",标志本回合 agent.reply_stream 收尾。
|
|
# ExceedMaxItersEvent(type=="EXCEED_MAX_ITERS")是另一种「本回合结束」收尾(内层 ReAct 撞 max_iters)。
|
|
_EVENT_REPLY_END = "REPLY_END"
|
|
_EVENT_EXCEED_MAX_ITERS = "EXCEED_MAX_ITERS"
|
|
# 本回合「结束」事件集合(读到其一即视为这轮 chat run 跑完,可评门)。
|
|
_TURN_END_EVENTS = frozenset({_EVENT_REPLY_END, _EVENT_EXCEED_MAX_ITERS})
|
|
|
|
|
|
def _infer_role(rel: str) -> str:
|
|
"""据工程内相对路径推断 A3 fileTree[].role(与 run.load_fixture_scaffold 的 _role 同口径)。
|
|
|
|
服务态控制面据 on-disk 文件重建 source_project,需要给每个文件标 role(交付契约 fileTree[].role)。
|
|
与 fixture 脚手架同一套推断规则,保证重建出的 fileTree role 与生成期一致(无 split-brain)。
|
|
"""
|
|
if rel == "src/main.js":
|
|
return "entry"
|
|
if "/scenes/" in rel:
|
|
return "scene"
|
|
if "/systems/" in rel or rel == "src/game-core.js":
|
|
return "system"
|
|
if "/data/" in rel:
|
|
return "config"
|
|
if "/util/" in rel or rel.endswith("seeded-random.js"):
|
|
return "lib"
|
|
if "/assets/" in rel:
|
|
return "asset-manifest"
|
|
return "other"
|
|
|
|
|
|
def _collect_on_disk_project(game_id: str) -> dict | None:
|
|
"""据 on-disk 工程目录(run._workdir/<game_id>)重建源工程交付形状(七要素)+ 文件全文清单。
|
|
|
|
【为什么由控制面据 on-disk 重建,而不取内存 session.finished】服务态九工具每回合工厂新建 Tier2Session,
|
|
内存态跨回合不保留(见文件头)。门绿后要落库的源工程,只能从九工具刚写盘的工程目录读回重建。
|
|
重建的形状对齐 worker.toolkit.finish 组装的 A3 源项目契约七要素(schemaVersion/projectType/fileTree/
|
|
entry/buildProfile/depLock/contentHash/addressing),保证落库口径与 CLI finish 一致(F3 无漂移)。
|
|
|
|
Returns:
|
|
{"source_project": {七要素}, "file_list": [{path, content}]} 或 None(目录不存在/无源文件)。
|
|
"""
|
|
# 惰性 import:run / toolkit 顶层牵出 agentscope(6c6g 未装),只在真要落库时才 import。
|
|
from worker import run as _run # noqa: PLC0415
|
|
from worker import toolkit as _toolkit # noqa: PLC0415
|
|
|
|
wd = _run._workdir(game_id)
|
|
if not wd.exists():
|
|
return None
|
|
|
|
# 收集 on-disk 源文件(排除构建产物 bundle / evidence / play-spec / index.html 这些非源文件)。
|
|
# 只收 src/ 与 data/ 下的源(对齐 toolkit 累积的工程文件树语义),其余是 build/run_gates 副产物。
|
|
file_list: list[dict] = []
|
|
for p in sorted(wd.rglob("*")):
|
|
if not p.is_file():
|
|
continue
|
|
rel = str(p.relative_to(wd)).replace("\\", "/")
|
|
# 跳过构建 / 真玩副产物(非交付源文件)。
|
|
if rel in ("bundle.iife.js", "play-spec.json", "index.html"):
|
|
continue
|
|
if rel.startswith("evidence/") or "/node_modules/" in ("/" + rel):
|
|
continue
|
|
try:
|
|
file_list.append({"path": rel, "content": p.read_text(encoding="utf-8")})
|
|
except Exception as e: # noqa: BLE001 —— 单文件读失败记日志、不中断重建其余文件
|
|
print(f"[tier2-control] ⚠ 重建源工程读文件失败 {rel}: {type(e).__name__}: {e}", flush=True)
|
|
|
|
if not file_list:
|
|
return None
|
|
|
|
# 组装 A3 源项目契约七要素(与 toolkit.finish:344-357 同口径;contentHash 复用 toolkit._content_hash)。
|
|
file_tree = [{"path": f["path"], "role": _infer_role(f["path"])} for f in file_list]
|
|
chash = _toolkit._content_hash(file_list)
|
|
source_project = {
|
|
"schemaVersion": "tier2-1.0",
|
|
"projectType": "tier2-phaser",
|
|
"fileTree": file_tree,
|
|
"entry": "src/main.js",
|
|
"buildProfile": {
|
|
"bundler": "esbuild", "format": "iife",
|
|
"globalName": _run.DEFAULT_GLOBAL_NAME, "minify": True, "target": "es2019",
|
|
},
|
|
# depLock.phaser:据 on-disk 无法可靠读出精确版本,用 toolkit 默认值(与 finish 同源)。
|
|
"depLock": {"phaser": "3.80.1"},
|
|
"contentHash": chash,
|
|
"addressing": {"store": "mysql+oss", "fetchById": "contentHash"},
|
|
}
|
|
return {"source_project": source_project, "file_list": file_list}
|
|
|
|
|
|
async def _wait_for_turn_end(base_url: str, agent_id: str, session_id: str, *,
|
|
user_id: str, timeout_s: float, idle_timeout_s: float,
|
|
trace_path: Any = None, attempt: int = 0) -> dict:
|
|
"""消费 SSE GET /sessions/{session_id}/stream,等本回合 chat run 结束(读到 REPLY_END / EXCEED_MAX_ITERS)。
|
|
|
|
源码核验(agentscope/app/_router/_session.py:425-534):
|
|
- 端点 = GET /sessions/{session_id}/stream,**需带 query agent_id**(:432);
|
|
- 返回 text/event-stream,帧形如 `data: {json}\\n\\n`,空闲每 30s 发心跳注释 `:\\n\\n`(:446/522);
|
|
- 流先 replay 当前 run 的缓冲事件,再 live 订阅;订阅是长连(不随单回合结束而关,只随 bus 关闭结束),
|
|
所以本函数靠「读到本回合的 REPLY_END/EXCEED_MAX_ITERS」自行收尾,不等服务端关流。
|
|
|
|
兜底纪律(绝不卡死):
|
|
- 总超时 timeout_s:超时即返回 {ended: False, reason: 'total_timeout'},由上层据 on-disk 继续评门;
|
|
- 空闲超时 idle_timeout_s:连续这么久没收到任何 data 帧(只有心跳/无字节)→ 视为断流,返回
|
|
{ended: False, reason: 'idle_timeout'},同样据 on-disk 继续评门(SSE 断流不致命)。
|
|
|
|
Returns:
|
|
{ended: bool, reason: str, endEvent: dict|None}。ended=True 表示读到本回合结束事件。
|
|
"""
|
|
import httpx # noqa: PLC0415 —— 惰性 import 红线(6c6g 未必装 httpx)
|
|
|
|
headers = {"X-User-Id": user_id} if user_id else {}
|
|
url = f"{base_url}/sessions/{session_id}/stream"
|
|
params = {"agent_id": agent_id}
|
|
t_start = time.perf_counter()
|
|
last_event_at = t_start
|
|
print(f"[tier2-control] 订阅 SSE 等本回合结束: session={session_id} agent={agent_id} "
|
|
f"(总超时 {timeout_s:.0f}s / 空闲超时 {idle_timeout_s:.0f}s)", flush=True)
|
|
try:
|
|
# read=idle_timeout_s(2026-06-24 cp-smoke-008 实证修复):此前 read=None,SSE 流【完全静默】时
|
|
# (008:服务端 3 次 redis 超时打断了事件发布 → 既无 data 也无心跳)aiter_lines 永久 await,
|
|
# 而总/空闲超时检查写在 line 循环【体内】、无 line 进来就永不触发 → 控制面卡死 60min(实证)。
|
|
# 把 read 设成 idle_timeout_s:静默超过它即 httpx.ReadTimeout → 下方 except 兜住 → 据 on-disk 评门继续。
|
|
# idle_timeout_s=300s 远大于 agent 内部 run_gates(chrome ~60s)的正常静默间隙,不会误杀正常回合。
|
|
timeout_cfg = httpx.Timeout(connect=15.0, read=idle_timeout_s, write=15.0, pool=15.0)
|
|
async with httpx.AsyncClient(timeout=timeout_cfg) as http:
|
|
async with http.stream("GET", url, params=params, headers=headers) as resp:
|
|
if resp.status_code >= 300:
|
|
body = (await resp.aread())[:300]
|
|
print(f"[tier2-control] ⚠ SSE 订阅非 2xx: {resp.status_code}: {body!r}", flush=True)
|
|
return {"ended": False, "reason": f"sse_http_{resp.status_code}", "endEvent": None}
|
|
async for line in resp.aiter_lines():
|
|
now = time.perf_counter()
|
|
# 总超时:本回合等太久,据 on-disk 继续(不卡死)。
|
|
if now - t_start > timeout_s:
|
|
print(f"[tier2-control] ⚠ SSE 等本回合结束总超时(>{timeout_s:.0f}s),"
|
|
"据 on-disk 继续评门。", flush=True)
|
|
return {"ended": False, "reason": "total_timeout", "endEvent": None}
|
|
if not line:
|
|
# 空行(SSE 帧分隔)——不更新空闲计时(只有真 data 帧才算「有进展」)。
|
|
if now - last_event_at > idle_timeout_s:
|
|
print(f"[tier2-control] ⚠ SSE 空闲超时(>{idle_timeout_s:.0f}s 无 data 帧),"
|
|
"疑断流,据 on-disk 继续评门。", flush=True)
|
|
return {"ended": False, "reason": "idle_timeout", "endEvent": None}
|
|
continue
|
|
if line.startswith(":"):
|
|
# 心跳注释帧(:\n\n);连接活着但本回合还没结束,据空闲超时判断是否断流。
|
|
if now - last_event_at > idle_timeout_s:
|
|
print(f"[tier2-control] ⚠ SSE 只收到心跳、空闲超时(>{idle_timeout_s:.0f}s),"
|
|
"据 on-disk 继续评门。", flush=True)
|
|
return {"ended": False, "reason": "idle_timeout", "endEvent": None}
|
|
continue
|
|
if not line.startswith("data:"):
|
|
continue
|
|
# 收到一个 data 帧:解析事件,刷新空闲计时。
|
|
last_event_at = now
|
|
payload = line[len("data:"):].strip()
|
|
if not payload:
|
|
continue
|
|
try:
|
|
event = json.loads(payload)
|
|
except Exception as e: # noqa: BLE001 —— 单帧解析失败记日志、不中断订阅
|
|
print(f"[tier2-control] ⚠ SSE 帧解析失败(已跳过): {type(e).__name__}: {e}", flush=True)
|
|
continue
|
|
etype = event.get("type")
|
|
# C1 观测:把结构性事件(跳过 *_DELTA 逐 token 噪声)记进 per-run trace 时间线。
|
|
if etype and not etype.endswith("_DELTA"):
|
|
_trace_sse_event(trace_path, event, now - t_start, attempt)
|
|
# 只认本回合的结束事件(REPLY_END / EXCEED_MAX_ITERS);其余事件(模型流 / 工具调用)略过。
|
|
if etype in _TURN_END_EVENTS:
|
|
print(f"[tier2-control] SSE 读到本回合结束事件 type={etype}"
|
|
f"(reply_id={event.get('reply_id')});评门。", flush=True)
|
|
return {"ended": True, "reason": etype, "endEvent": event}
|
|
# 流自然结束(bus 关闭等)——本回合可能已收尾,据 on-disk 评门。
|
|
print("[tier2-control] ⚠ SSE 流结束但未读到本回合结束事件,据 on-disk 继续评门。", flush=True)
|
|
return {"ended": False, "reason": "stream_closed", "endEvent": None}
|
|
except Exception as e: # noqa: BLE001 —— SSE 任何异常都 best-effort:不卡死,据 on-disk 继续评门
|
|
print(f"[tier2-control] ⚠ SSE 订阅异常(据 on-disk 继续评门): {type(e).__name__}: {e}", flush=True)
|
|
return {"ended": False, "reason": f"sse_error:{type(e).__name__}", "endEvent": None}
|
|
|
|
|
|
# ── C1 轻量观测:把 SSE 事件流记成每局可读 trace(客户端侧,零服务改动 / 零 Studio 依赖)──
|
|
# 迭代痛点 = 看不见 run 在哪卡住。控制面本就逐帧消费 SSE,顺手把【结构性事件】(跳过 *_DELTA 逐 token 噪声)
|
|
# 记成一份 per-run JSONL 时间线:模型调用 / 工具调用(带工具名)/ 工具结果 / 推理 / reply 边界 / 卡死信号
|
|
# (REQUIRE_USER_CONFIRM / REQUIRE_EXTERNAL_EXECUTION)。落在工程 workdir/control-plane-trace.jsonl,
|
|
# 跨 attempt 追加 = 整局多轮时间线。best-effort:写失败只告警、绝不中断主链。这是「完整的 C」里 C1 观测的
|
|
# 轻量兑现(无需起 Node Studio;Studio/OTLP 那条 observability.studio_sink 仍在、是将来全量 UI 的可选路)。
|
|
def _compact_sse_event(event: dict) -> dict:
|
|
"""把一个 SSE 事件压成 trace 行的精简形(只留有意义字段,长文本截断;绝不抛)。"""
|
|
etype = event.get("type")
|
|
out: dict = {"type": etype, "reply_id": event.get("reply_id")}
|
|
# 工具调用:记工具名 + call-id + 结果态(看 agent 调了哪些工具、按什么顺序、成没成——卡在哪一目了然)。
|
|
# 2.0.2 真实字段名(源码核实 event/_event.py):TOOL_CALL_START 带 tool_call_name + tool_call_id;
|
|
# TOOL_CALL_END / TOOL_RESULT_END 带 tool_call_id(+ state=结果态)。此前误取 tool_name/name/id → 抓不到名。
|
|
for k in ("tool_call_name", "tool_call_id", "tool_name", "name", "state"):
|
|
if event.get(k):
|
|
out[k] = event.get(k)
|
|
# 截断可能的长文本字段(text/delta/output/content 的字符串形),只留前 200 字。
|
|
for k in ("text", "output", "content", "reason", "error"):
|
|
v = event.get(k)
|
|
if isinstance(v, str) and v:
|
|
out[k] = v[:200]
|
|
return out
|
|
|
|
|
|
def _trace_sse_event(trace_path: Any, event: dict, elapsed_s: float, attempt: int) -> None:
|
|
"""把一个结构性 SSE 事件追加进 per-run trace JSONL(best-effort,绝不抛)。"""
|
|
if not trace_path:
|
|
return
|
|
try:
|
|
import json as _json # noqa: PLC0415
|
|
import os as _os # noqa: PLC0415
|
|
# 工程 workdir 尚未建(scaffold_init 之前的早期事件)→ 静默跳过,不抢先建目录(免干扰 scaffold 的
|
|
# "已存在不覆盖" 逻辑),也免刷 FileNotFoundError(008 实证:首轮 scaffold 前会有这类早期事件)。
|
|
if not _os.path.isdir(_os.path.dirname(trace_path)):
|
|
return
|
|
rec = {"t": round(elapsed_s, 1), "attempt": attempt, **_compact_sse_event(event)}
|
|
with open(trace_path, "a", encoding="utf-8") as f:
|
|
f.write(_json.dumps(rec, ensure_ascii=False) + "\n")
|
|
except Exception as e: # noqa: BLE001 —— 观测落盘失败绝不中断主链(C1 observe-only)
|
|
print(f"[tier2-control][trace] ⚠ trace 落盘失败(忽略):{type(e).__name__}: {e}", flush=True)
|
|
|
|
|
|
# ── 续跑指令(等价 CLI run_studio:454-458 的「停了但门没绿,按失败门继续修、门绿再 finish」)──
|
|
def _resume_feedback_text(verdict_feedback: str) -> str:
|
|
"""把失败门反馈包成续跑指令(口径对齐 studio.py:454-458)。"""
|
|
return (
|
|
"你刚才停下了,但验收门还没全绿——不要放弃。这是上次 run_gates 的失败门:\n"
|
|
f"{verdict_feedback}\n"
|
|
"请在循环里:据失败门 write_source 针对性修(数据表 schema 错就先 validate_datatable 看平台要的 key),"
|
|
"build→run_gates→read_verdict,直到门绿再 finish。一步步来,先修最关键的致命门。")
|
|
|
|
|
|
async def drive_generation(
|
|
base_url: str,
|
|
game_id: str,
|
|
brief: str,
|
|
play_spec: dict | None = None,
|
|
*,
|
|
max_resumes: int | None = None,
|
|
model_name: str = "MiniMax-M3",
|
|
writer_max_iters: int | None = None,
|
|
user_id: str = "tier2",
|
|
do_design: bool = True,
|
|
single_post: bool = True,
|
|
sse_turn_timeout_s: float = 1200.0,
|
|
sse_idle_timeout_s: float = 300.0,
|
|
) -> dict:
|
|
"""服务态有界 resume 控制面:驱动一款富游戏从启动到门绿落库(等价 CLI run_studio 的外层 resume)。
|
|
|
|
经 bootstrap 的现成 REST 原语 + run 的纯代码门驱动 service/app.py 的 create_app;**绝不重写 REST/生成逻辑**。
|
|
收敛判据全靠 run.run_gates 机器判门(零自评、绝不放松门),落库据 on-disk 源工程(不靠内存 session)。
|
|
|
|
Args:
|
|
base_url: 运行中的 Agent Service 根地址(如 http://100.64.0.7:8200)。
|
|
game_id: 调用方语义上的游戏标识(仅用于日志/返回;真实评门 game_id = 框架分配的 session_id,见下)。
|
|
brief: 一句话题面。
|
|
play_spec: 真玩驱动规格(透传 run.run_gates;None 走默认 business-sim driver)。
|
|
max_resumes: 外层 resume 上限;None → genconfig.get('iteration','max_resumes',6)(对齐 studio)。
|
|
model_name: M3 模型名(默认 MiniMax-M3,经 new-api 走 Anthropic 原生)。
|
|
writer_max_iters: 单写 ReAct 放开轮数;None → genconfig.get('iteration','writer_max_iters',40)。
|
|
user_id: 多租户用户标识(经 X-User-Id 头)。
|
|
single_post: True(默认,阶段一①)= 只发一次 kick,续修在 Service 端 on_reasoning middleware 内完成,
|
|
消费方等这一次回合真结束后读服务端已落 verdict 判落库,不再外层多轮 resume;False = 走现有外层
|
|
有界 resume 循环(fallback,续修 middleware spike 不成时回落)。
|
|
sse_turn_timeout_s: 单回合 SSE 等结束的总超时(兜底,绝不卡死)。
|
|
sse_idle_timeout_s: 单回合 SSE 空闲(无 data 帧)断流判定超时。
|
|
|
|
Returns:
|
|
{game_id, agent_id, session_id, finished(bool=门绿+落库), attempts, last_verdict,
|
|
store_addressing, wall_s, stopped_reason}。
|
|
"""
|
|
# 惰性 import:bootstrap(REST 原语,牵出 httpx/worker)/ run(门 + 落库,牵出 agentscope)/ genconfig /
|
|
# gate_judge(single_post 读服务端已落 verdict 后经它归一判门绿,与门线口径一致、不放松)。
|
|
from worker import genconfig, run as _run # noqa: PLC0415
|
|
from worker.gate_judge import judge_tier2_verdict # noqa: PLC0415
|
|
from . import bootstrap # noqa: PLC0415
|
|
|
|
t0 = time.perf_counter()
|
|
# 旋钮口径对齐 studio:max_resumes / writer_max_iters 缺省从 generation.yaml 读。
|
|
if max_resumes is None:
|
|
max_resumes = genconfig.get("iteration", "max_resumes", 6)
|
|
if writer_max_iters is None:
|
|
writer_max_iters = genconfig.get("iteration", "writer_max_iters", 40)
|
|
|
|
result: dict[str, Any] = {
|
|
"game_id": game_id, "agent_id": None, "session_id": None,
|
|
"finished": False, "attempts": 0, "last_verdict": None,
|
|
"store_addressing": None, "wall_s": 0.0, "stopped_reason": None,
|
|
}
|
|
|
|
# ── ① 启动:bootstrap.start_new_game(注册凭据→建 agent→建 session→发 kick)──
|
|
# 关键集成点:评门用的 game_id 必须 = 框架分配的 session_id(九工具据 session_id 当 game_id 管目录,
|
|
# 见 app.py:_tier2_tools_factory:150),否则 run_gates 评的不是九工具刚写的那个工程目录。
|
|
try:
|
|
started = await bootstrap.start_new_game(
|
|
base_url, brief, user_id=user_id, model_name=model_name,
|
|
writer_max_iters=writer_max_iters, do_design=do_design)
|
|
except Exception as e: # noqa: BLE001 —— 启动失败 best-effort 落结果、记日志,不静默吞
|
|
result["stopped_reason"] = f"start_failed:{type(e).__name__}: {e}"
|
|
result["wall_s"] = round(time.perf_counter() - t0, 1)
|
|
print(f"[tier2-control] game={game_id} 启动失败:{type(e).__name__}: {e}", flush=True)
|
|
return result
|
|
|
|
agent_id = started.get("agent_id")
|
|
session_id = started.get("session_id")
|
|
result["agent_id"] = agent_id
|
|
result["session_id"] = session_id
|
|
# 评门 game_id = 框架 session_id(与九工具工程目录一致)。
|
|
gate_game_id = session_id
|
|
print(f"[tier2-control] game={game_id} 启动: agent={agent_id} session={session_id} "
|
|
f"(评门 game_id=session_id;max_resumes={max_resumes})", flush=True)
|
|
# C1 观测:本局 SSE 事件 trace 落工程 workdir/control-plane-trace.jsonl(跨 attempt 追加,可读时间线)。
|
|
trace_path = None
|
|
try:
|
|
trace_path = str(_run._workdir(gate_game_id) / "control-plane-trace.jsonl")
|
|
print(f"[tier2-control] 观测 trace → {trace_path}", flush=True)
|
|
except Exception: # noqa: BLE001 —— 取 trace 路径失败不影响主链(观测 best-effort)
|
|
trace_path = None
|
|
if not session_id or not agent_id:
|
|
result["stopped_reason"] = "start_missing_ids(start_new_game 未回 agent_id/session_id)"
|
|
result["wall_s"] = round(time.perf_counter() - t0, 1)
|
|
return result
|
|
|
|
# ── ②' 单 POST 路(阶段一①):续修在 Service 端 on_reasoning middleware 内完成,消费方只发一次 kick、
|
|
# 等这一次(内部续修多轮)回合真结束、读服务端已落 verdict 判落库;不再外层多轮 resume。老循环留 fallback。──
|
|
if single_post:
|
|
# SSE 总超时须覆盖续修 N 次串行真门(评审 C-2):调用方未显式放大时按 budget.repair_wall_timeout_s 兜底,
|
|
# 与工厂 breaker 的 repair_wall_s 同源对齐;绝不用默认 1200s(否则续修中的合法局被误判 not_ended、丢产物)。
|
|
sse_timeout = max(sse_turn_timeout_s,
|
|
genconfig.get("budget", "repair_wall_timeout_s",
|
|
(max_resumes + 1) * (300 + 120) + 300))
|
|
turn = await _wait_for_turn_end(
|
|
base_url, agent_id, session_id, user_id=user_id,
|
|
timeout_s=sse_timeout, idle_timeout_s=sse_idle_timeout_s,
|
|
trace_path=trace_path, attempt=0)
|
|
result["attempts"] = 1
|
|
if not turn.get("ended"):
|
|
# 回合未真结束(SSE 总超时/断流):服务端可能仍在续修写盘,评中间态会误判 + 与写盘竞态 → 不评、不落库。
|
|
print(f"[tier2-control] game={game_id} single_post 回合未真结束"
|
|
f"(reason={turn.get('reason')}),不评中间态、不落库。", flush=True)
|
|
result["stopped_reason"] = f"single_post_turn_not_ended:{turn.get('reason')}"
|
|
result["wall_s"] = round(time.perf_counter() - t0, 1)
|
|
return result
|
|
# 回合真结束(REPLY_END/EXCEED_MAX_ITERS):读服务端续修已落的 verdict(不重跑门,评审 I5)。
|
|
v = _run.read_last_verdict(gate_game_id)
|
|
# gameId 带上(F-1):判失败时 judgment.feedback 才含经济数值证据(与门线 run_gates 返回口径一致)。
|
|
judgment = judge_tier2_verdict({"rc": 0, "verdict": v, "log": "", "gameId": gate_game_id})
|
|
result["last_verdict"] = v or None
|
|
if judgment.passed:
|
|
# 门绿:据 on-disk 源工程落库(逻辑同外层循环 c 分支)。
|
|
built = _collect_on_disk_project(gate_game_id)
|
|
if built:
|
|
addr = _run.persist_source_project(
|
|
gate_game_id, built["source_project"], built["file_list"], now_ts=time.time())
|
|
result["store_addressing"] = addr
|
|
result["finished"] = bool(addr)
|
|
result["stopped_reason"] = "gates_green_persisted" if addr else "gates_green_but_persist_failed"
|
|
else:
|
|
result["stopped_reason"] = "gates_green_but_no_on_disk_source"
|
|
else:
|
|
# 软停/续修耗尽放行的尽力产物:门没绿 → 只留 on-disk workdir、不入 store(创始人 2026-07-02:质量门守住)。
|
|
print(f"[tier2-control] game={game_id} single_post 门未绿(未过门={judgment.failed_gates}),"
|
|
"尽力产物留 on-disk workdir、不入库。", flush=True)
|
|
result["stopped_reason"] = "single_post_gates_not_green"
|
|
result["wall_s"] = round(time.perf_counter() - t0, 1)
|
|
print(f"[tier2-control] game={game_id} single_post 结束: finished={result['finished']} "
|
|
f"reason={result['stopped_reason']} wall={result['wall_s']}s", flush=True)
|
|
return result
|
|
|
|
last_verdict: dict | None = None
|
|
# ── ② 外层有界 resume 循环(fallback:single_post=False;续修 middleware spike 不成时回落这条,设计 §6)──
|
|
for attempt in range(max_resumes + 1):
|
|
result["attempts"] = attempt + 1
|
|
# a. 等本回合 chat run 结束(SSE;超时/断流据 on-disk 继续,绝不卡死)。
|
|
await _wait_for_turn_end(
|
|
base_url, agent_id, session_id, user_id=user_id,
|
|
timeout_s=sse_turn_timeout_s, idle_timeout_s=sse_idle_timeout_s,
|
|
trace_path=trace_path, attempt=attempt)
|
|
|
|
# b. 独立评门:run.run_gates 纯代码判 on-disk 源工程(零 LLM、确定性、绝不自评翻绿)。
|
|
try:
|
|
gate = _run.run_gates(gate_game_id, play_spec)
|
|
except Exception as e: # noqa: BLE001 —— 评门异常 best-effort:记日志,带空 verdict 进续修分支
|
|
print(f"[tier2-control] game={game_id} attempt={attempt} run_gates 异常:"
|
|
f"{type(e).__name__}: {e}", flush=True)
|
|
gate = {"rc": 1, "verdict": None, "log": f"run_gates 异常:{type(e).__name__}: {e}"}
|
|
v = gate.get("verdict") or {}
|
|
last_verdict = v or last_verdict
|
|
result["last_verdict"] = last_verdict
|
|
|
|
# c. 门绿判定(decision==accept 且 L1.passed)。
|
|
l1 = ((v.get("layerResults") or {}).get("L1") or {})
|
|
if v.get("decision") == "accept" and l1.get("passed"):
|
|
print(f"[tier2-control] game={game_id} attempt={attempt} 验收门全绿,据 on-disk 落库。", flush=True)
|
|
# 据 on-disk 源工程落库(不依赖内存 session.finished)。
|
|
built = _collect_on_disk_project(gate_game_id)
|
|
if built:
|
|
addr = _run.persist_source_project(
|
|
gate_game_id, built["source_project"], built["file_list"], now_ts=time.time())
|
|
result["store_addressing"] = addr
|
|
result["finished"] = bool(addr) # 落库成功才算真 finished(门绿 + 落库)
|
|
if not addr:
|
|
print(f"[tier2-control] game={game_id} ⚠ 门绿但落库失败(产物仍在 GEN_DIR workdir);"
|
|
"stopped_reason=persist_failed。", flush=True)
|
|
result["stopped_reason"] = "gates_green_but_persist_failed"
|
|
else:
|
|
result["stopped_reason"] = "gates_green_persisted"
|
|
else:
|
|
# 门绿却读不到 on-disk 源工程(异常态)——诚实标注,不伪造落库。
|
|
print(f"[tier2-control] game={game_id} ⚠ 门绿但 on-disk 工程目录读不到源文件,无法落库。",
|
|
flush=True)
|
|
result["stopped_reason"] = "gates_green_but_no_on_disk_source"
|
|
break
|
|
|
|
# d. 预算耗尽:已是最后一次 attempt 仍未门绿 → 停。
|
|
if attempt >= max_resumes:
|
|
print(f"[tier2-control] game={game_id} resume 预算耗尽({max_resumes} 次)仍未门绿,停。",
|
|
flush=True)
|
|
result["stopped_reason"] = "budget_exhausted"
|
|
break
|
|
|
|
# e. 门没绿 → 带 verdict 失败反馈 resume 续修(头号自纠机制;等价 studio.py:451-458;
|
|
# 带 game_id 使经济门失败附三个数数值证据 + H 门附断言清单,F-1 反馈契约)。
|
|
fb = _run.verdict_feedback(v, gate.get("log") or "", game_id=gate_game_id) if v else \
|
|
("尚无 verdict(疑真玩未产出 / harness 未就位)。" + (gate.get("log", "")[:400]))
|
|
print(f"[tier2-control] game={game_id} attempt={attempt} 门未绿"
|
|
f"(decision={v.get('decision')}),带反馈 resume 续修。", flush=True)
|
|
try:
|
|
await bootstrap.resume_session(
|
|
base_url, agent_id, session_id, user_id=user_id,
|
|
feedback=_resume_feedback_text(fb))
|
|
except Exception as e: # noqa: BLE001 —— resume 抛 = 熔断:best-effort 落部分结果返回,绝不静默吞
|
|
print(f"[tier2-control] game={game_id} attempt={attempt} resume 异常(熔断):"
|
|
f"{type(e).__name__}: {e}", flush=True)
|
|
result["stopped_reason"] = f"resume_failed:{type(e).__name__}: {e}"
|
|
break
|
|
|
|
if result["stopped_reason"] is None:
|
|
# 循环正常走完未命中任何分支(理论不达;兜底标注)。
|
|
result["stopped_reason"] = "loop_ended"
|
|
result["wall_s"] = round(time.perf_counter() - t0, 1)
|
|
print(f"[tier2-control] game={game_id} 结束: finished={result['finished']} "
|
|
f"attempts={result['attempts']} reason={result['stopped_reason']} "
|
|
f"wall={result['wall_s']}s", flush=True)
|
|
return result
|
|
|
|
|
|
def main(argv: list[str] | None = None) -> int:
|
|
"""CLI 入口:python -m service.control_plane <game_id> --base-url ... --brief ...。
|
|
|
|
惰性 import asyncio / argparse(顶层零重依赖红线;此处也只用标准库)。真跑在 mini-desktop。
|
|
"""
|
|
import argparse # noqa: PLC0415
|
|
import asyncio # noqa: PLC0415
|
|
|
|
parser = argparse.ArgumentParser(
|
|
description="tier2 服务态有界 resume 控制面:驱动一款富游戏到门绿落库。")
|
|
parser.add_argument("game_id", help="游戏标识(用于日志/返回;真实评门用框架分配的 session_id)。")
|
|
parser.add_argument("--base-url", required=True,
|
|
help="运行中的 Agent Service 根地址(如 http://100.64.0.7:8200)。")
|
|
parser.add_argument("--brief", required=True, help="一句话题面。")
|
|
parser.add_argument("--max-resumes", type=int, default=None,
|
|
help="外层 resume 上限(默认从 generation.yaml iteration.max_resumes 读)。")
|
|
parser.add_argument("--writer-max-iters", type=int, default=None,
|
|
help="单写 ReAct 放开轮数(默认从 generation.yaml iteration.writer_max_iters 读)。")
|
|
parser.add_argument("--model-name", default="MiniMax-M3", help="模型名(默认 MiniMax-M3)。")
|
|
parser.add_argument("--user-id", default="tier2", help="多租户用户标识(默认 tier2)。")
|
|
parser.add_argument("--no-design", action="store_true",
|
|
help="跳过阶段 1 工作室设计(writer 拿空 design_text;调试/对照用)。默认跑设计。")
|
|
args = parser.parse_args(argv)
|
|
|
|
res = asyncio.run(drive_generation(
|
|
args.base_url, args.game_id, args.brief,
|
|
max_resumes=args.max_resumes, writer_max_iters=args.writer_max_iters,
|
|
model_name=args.model_name, user_id=args.user_id, do_design=not args.no_design))
|
|
# 结构化结果打到 stdout(可被上层脚本捕获)。
|
|
print(json.dumps(res, ensure_ascii=False, indent=2), flush=True)
|
|
# 退出码:门绿落库成功 = 0,否则 1(供编排脚本判成败)。
|
|
return 0 if res.get("finished") else 1
|
|
|
|
|
|
if __name__ == "__main__":
|
|
import sys # noqa: PLC0415
|
|
|
|
sys.exit(main())
|