docs(plan): plan003 Plan A gen-done 执行计划(用对 M3→≥80%→两步 flip)
Plan A = gen-done 生死门 = 003-U1(用对 agentic M3) + 003-U2(两步 flip)。 已过双评审(Opus 对抗式+可行性,代码核证)整改 5 P0 + 7 P1: - Anthropic 多 Generation 响应形(block0=thinking,须遍历 getResults) - checkpoint 序列化潜伏坑(e2e harness checkpoint-off 隐藏,staging 才炸) - 独立性纠偏(量测执行 Plan B gd-runtime/build-from-source → KTD8 运行时钉) - ≥80% 机制门口径(对资产质量盲)+ 并发共享面审计 + measure-first U-C 闸控 对齐 Plan B U1(db14845d,asset-orphan 已闭合)为收敛运行时基底。 Codex(codex-rescue) 未出可用评审 → §6.8 Opus 单评回退。 Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
This commit is contained in:
parent
079778fc03
commit
b6cadd8394
357
docs/plans/2026-06-19-002-feat-gendone-用对M3生成生死门-plan.md
Normal file
357
docs/plans/2026-06-19-002-feat-gendone-用对M3生成生死门-plan.md
Normal file
@ -0,0 +1,357 @@
|
||||
---
|
||||
title: "feat: plan003 Plan A — gen-done 生成生死门(用对 agentic 的 M3 → ≥80% → 两步 flip)"
|
||||
status: active
|
||||
date: 2026-06-19
|
||||
type: feat
|
||||
origin: docs/plans/2026-06-18-003-feat-gen-quality-converge-and-m0-integration-plan.md(003-U1/U2 骨架)+ docs/brainstorms/2026-06-18-生成失败根因分析-requirements.md(v3「用对 M3」方向 SoT)+ docs/brainstorms/2026-06-19-plan003-phase1验收闭环-requirements.md(R1/R2/T1/T2 done 判据)
|
||||
review: "已过双评审(AGENTS.md §6 条款 8,2026-06-19):Opus 对抗式 + Opus 可行性双评(代码核证);Codex(codex-rescue) 11min 未出可用评审 → §6.8 Opus 单评回退。整改 5 P0 + 7 P1 + 多 P2(见「双评审整改注记」)。ce-work 前可补跑 Codex 作 belt-and-suspenders。"
|
||||
---
|
||||
|
||||
# feat: plan003 Plan A — gen-done 生成生死门
|
||||
|
||||
> **双线拆分**:本档 = **Plan A(gen-done)**;Plan B(core-done)= `docs/plans/2026-06-19-001-feat-coredone-验收闭环执行-plan.md`,另一 session(worktree `feat/mac-planB-coredone`)并行。
|
||||
>
|
||||
> **⚠️ 边界的真相(评审 P0-1 纠偏)**:Plan A 与 Plan B **代码作者面不相交**(Plan A **不编辑** `gd-runtime.js`/`build-from-source.mjs`,Plan B **不编辑** SAA java / worker python)。**但"作者面不相交 ≠ 量测面不相交"**:gen-done 的 ≥80% 量测**执行期实跑** Plan B 的两文件——validate 节点 `SaaGenNodes.java:283` 以子进程跑 `node src/host/build-from-source.mjs`,装配产物 `build-from-source.mjs:34` 硬 import `../../../src/host/gd-runtime.js`,九门真玩(F_wiring/H_progress/I_control)每跑必执行 gd-runtime。**故 gen-done 的指标是 gd-runtime/build-from-source 行为的函数**。→ 见 KTD 8(运行时版本钉 + Plan B 冻结协调)。
|
||||
>
|
||||
> **铁律边界(执行期不可越)**:① **绝不编辑** `gd-runtime.js`/`build-from-source.mjs`(Plan B 域);② `remaining` 语义修正只走 `SaaPrompts` 的 `DESIGN_SYSTEM` java 范例;③ mini-desktop e2e 跑前确认无 Plan B 占用(端口 4320/9222 单实例 flock);④ **收敛窗口期协调 Plan B 冻结这两文件**(每轮量测前钉其 content-hash,变更=重置基线,见 KTD 8)。
|
||||
>
|
||||
> **单向耦合**:Plan B 的 Phase 2/3 game-quality 门**消费** gen-done 产出(末段、单向)。
|
||||
>
|
||||
> **编号命名空间**:本档 unit `U1…U7` = Plan A 局部命名空间;对应 plan003 `003-U1`(用对 M3)+ `003-U2`(两步 flip)。喂 ce-work / 跨档引用带 `PlanA-Ux`。
|
||||
|
||||
---
|
||||
|
||||
## 双评审整改注记(2026-06-19 · Opus 对抗式 + 可行性,代码核证)
|
||||
|
||||
> 整改 5 P0 + 7 P1(全 IN-DOCUMENT 落本档)+ P2/确认。逐条可回溯到下文章节。
|
||||
|
||||
- **P0-1(独立性主张过宽)**:~~"gen-done 不依赖 Plan B"~~ → 改"**作者面不相交、量测面执行 Plan B 两文件**";增 KTD 8 运行时 content-hash 钉 + 收敛窗口 Plan B 冻结。(见 双线拆分 note / KTD 8 / U5 / R-10)
|
||||
- **P0-2(≥80% 是机制门、对资产质量盲)**:~~`GAMEDEF_KEYS` strip 掉 assets、`drawEntity` 仅图元~~ **【2026-06-19 更新:Plan B U1(`db14845d`,feat/mac-planB-coredone)已闭合 asset-orphan——`gd-runtime` 现支持 `shape:'sprite'`→`drawSprite`→`drawImage`(受控面 host 预载),assets 走 `sourceProject` 顶层抽取(非 `GAMEDEF_KEYS` strip)。"运行时不能渲染资产"已过期】**。但九门仍只 `D_render` 验**非空画面**、**不验资产对位正确性**,故"**≥80% 仅证机制、不验资产/好玩**"仍成立(资产正确性归 core-done R9/R11);U7 的 asset-orphan 烟测**现已可满足**(运行时有 sprite 路)。(见 R1 / KTD 4 / 范围边界 / U7)
|
||||
- **P0-3(并发共享面 > 端口)**:`serve-and-play.sh:44` 服整 runtime 根、`_shared/` 模板共读、`/tmp/wg1-*-$PORT.log` 可撞、`lsof|kill` 净场 port-keyed → 结构性竞态(非时序)。U3 加全共享面审计 + 同款×N 结构隔离断言;K=1 为记录基线权威。(见 KTD 3 / U3)
|
||||
- **P0-4(Anthropic 多 Generation 响应形,feasibility)**:`AnthropicChatModel` 每 content block 一 `AssistantMessage`,thinking 开时 **block0=thinking** → `getResult().getText()` 取到推理文、抽取必败。U1/U2 改**遍历 `resp.getResults()`** 重组有序块、答案取 text-role Generation;历史项 = 每轮 `List<AssistantMessage>`(thinking 先序)。(见 KTD 2 / U1 / U2)
|
||||
- **P0-5(checkpoint 序列化潜伏坑,feasibility)**:e2e harness `saaCheckpointEnabled=false`+dataSource=null → **MysqlSaver 序列化路从不被量测触发**;staging 默认 `saaCheckpointEnabled=true` → `List<AssistantMessage>` 若 Jackson 往返丢元素类型,**绿门后在生产炸**。U2 加 checkpoint-ON 序列化往返测试;U7 隔离验须 checkpoint 开。(见 KTD 2 / U2 / U7 / R-9)
|
||||
- **P1-1(measure-first 隐藏晚长杆)**:U-C 是最难/最未证单元,若闸控到晚期才建=时间压力下建最不确定单元。增 U5 并行 U-C 廉价可行性 spike + 具体触发线。(见 KTD 5 / U5 / U6 / 待确认)
|
||||
- **P1-2(≥80% 标的未验 + brief Goodhart)**:80% 是 iife 路继承数、gamedef 真基线 n=0;brief 集自控且偏易(`SaaFullGraphE2eTest.java:90` 注释自陈"偏向 harness 支持品类")。增 基线定标复核 + brief 集冻结/钉 hash + **留出集**(final R1 用未调优集)。(见 R1 / U3 / U5 / 待确认)
|
||||
- **P1-3(thinking 历史无界增长 vs 保真张力)**:append 全 thinking × 8 救场轮 → token 爆 + 每轮复发计费;窗口截断又毁 thinking/顺序。增 thinking-aware 有界截断策略 + max-repair 深度(8)测试 + 并发隔离测试。(见 KTD 2 / U2)
|
||||
- **P1-4(KeyStrategy 全量 ReplaceStrategy,feasibility)**:`SaaStudioGraph` keyFactory 对 `ALL_KEYS` 全赋 `ReplaceStrategy`;"append 策略"须特例化或**节点内 read-append-write**(对齐 `appendAttempt`)。U2 选后者。(见 U2)
|
||||
- **P1-5(并发改动 > Semaphore(K),feasibility)**:`newSingleThreadExecutor` 须同改有界池;端口注入须 per-job 替 `buildInputs` 常量、`finally` 释放防泄漏。(见 KTD 3 / U3)
|
||||
- **P1-6(flip 须同暴露 saa-model-protocol,feasibility P2-4)**:U7 只暴露 dispatcher+sourceMode → staging 仍跑 openai 默认协议 ≠ U5 量测的 anthropic。U7 加暴露 `saa-model-protocol` + R1 证据记在用协议(量==发)。(见 U7 / R2)
|
||||
- **P1-7(T1 回退是 flip 硬前置,非可延 TODO)**:回退路效力依赖 Plan B 域代码;T1 失败 → step-2 flip **阻塞**(跨 Plan 关键依赖)。(见 U7 / R-fallback)
|
||||
- **P2**:yaml 合并 config-presence 断言(U7);client 侧 `budget<maxTokens` fail-fast(U1);确认无 `spring-ai-autoconfigure-model-anthropic` 传递(U1);spike 实证 1.1.2 多 Generation 形(U1);`llm-base` 仅需覆盖才加 yaml(去噪,U7)。
|
||||
|
||||
---
|
||||
|
||||
## 摘要
|
||||
|
||||
把 SAA **真结构化(gameDefinition)生成路**从「用错 M3 + 历史不留 + 救场从头重生成 + 整图串行不可量」推到「**用对 agentic 的 M3(Anthropic 协议 + thinking 全保留 + 连续对话救场)+ 可量产线门 ≥80%(机制门口径)+ 默认产线 flip(含回退验过)**」。
|
||||
|
||||
**第一性纠偏(代码核证)**:现状跨 Java/Python 三层**完全一致**——每轮全新 `[system,user]`、`getText()` 丢弃 thinking、九门走外部 `bash` 子进程当末端反馈文本、救场命令"重出完整模块不要 diff"。四项目标能力(Anthropic 协议 / 历史保留 / 门作工具 / staging dispatcher env)**全为净新建,无半成品可改**。可行性已实证:Anthropic+thinking+工具循环在现 `spring-ai-bom:1.1.2` 上 **FEASIBLE**(加一 bare 依赖、零 bean 冲突、阻塞调用),唯响应是**多 Generation 形**(block0=thinking,须遍历 `getResults()`,见 KTD 2)。
|
||||
|
||||
**执行姿势 = 先量真基线、数据驱动收敛**:用对 M3(U-A/U-B)+ 8/8 driver 上下文落地后,跑**有界并发**(K=正确性参数、非提速旋钮)n≥30 真基线(mini-desktop 权威 / lili-mac 快走查),用 **Workflow-Opus 全量根因分析**产出排序优化项,改→重测,循环到 ≥80%;**模型驱动工具循环(U-C)= measure-first 闸控 + 并行廉价 spike 去险**。达标后两步 flip(隔离验→切默认→回退验,含 saa-model-protocol 同切)。
|
||||
|
||||
> **≥80% 口径警示(评审 P0-2)**:本门是**机制门**(九门验引擎调用/进度/控制),**对资产渲染与"好玩"盲**(gamedef 装配 strip 掉 assets、运行时仅画图元)。≥80% 仅证"机制可玩",**不证产品级质量**——资产/玩法完整性由 core-done(Plan B R9/R11,003-U5)把守。flip 不可解读为"gamedef 已达产品质量"。
|
||||
|
||||
> 上游权威:plan003 `003-U1/U2` + 根因分析 v3 + phase1 验收闭环 R1/R2/T1/T2。本档 = 其 **HOW 执行层**,质量域续做、不另起平行 prompt。
|
||||
|
||||
---
|
||||
|
||||
## 问题背景(代码核证 · 含发现)
|
||||
|
||||
> 行号锚点为 2026-06-19 dev/2.0.0 rebase 后快照;执行期以类/方法名为准。
|
||||
|
||||
- **三层用法一致地"用错 M3"**:Java `SaaStudioNodes.callAndRecord`(~235-248)`getText()` 丢 thinking、每轮全新 `[SystemMessage,UserMessage]` 零历史;`SaaPrompts.buildMessages`(~348-356)两元 + "重出完整模块不要 diff";`repairNode`(~764-783)不回传上一版源(a3 未做);Python `_client.chat` 显式 `thinking:disabled`(用错的反面)。
|
||||
- **🔴 发现①:M3 双协议**。new-api `/anthropic` `/v1/messages` 返原生 `['thinking','text']` 块 + tool_use(探针证);OpenAI 默认路 thinking 内联撑爆截断。
|
||||
- **🔴 发现②:e2e 门当前量错路**。`SaaFullGraphE2eTest`(~203-211)构造 props **未 set `saaSourceMode`** → 默认 `factory`(iife),故现 60%(n=5)**是 iife 旧路、非 gameDefinition**。dispatcher=saa 单切仍走 iife,须同时 `saaSourceMode=gamedef`。
|
||||
- **🔴 发现③:整图串行 + 并发会污染门 + 共享面 > 端口**。`SaaGraphDispatcher`:`newSingleThreadExecutor`(~101)+ `Semaphore(1)`(~113)包整条 `graph.invoke`(~268 acquire→~275 invoke→~296 release)。串行因 = 九门 `serve-and-play.sh` 绑固定端口 4320/9222。play 节点(`SaaGenNodes` ~454-456)**已从 state 读 port/cdpPort**、`buildInputs`(~399-400)注入(现硬编码 4320/9222)。**但并发共享面 > 端口**(评审 P0-3):`serve-and-play.sh:44` 服整 runtime 根、`_shared/` 模板共读、`/tmp/wg1-*-$PORT.log` 可撞、`lsof|kill` 净场 port-keyed(同端口误杀兄弟)。
|
||||
- **🔴 发现④:gen-done 量测执行 Plan B 两文件**(评审 P0-1):`SaaGenNodes.java:283` 子进程跑 `build-from-source.mjs`;产物 `build-from-source.mjs:34` 硬 import `gd-runtime.js`;`build-from-source.mjs:32` `GAMEDEF_KEYS=['entities','components','behaviors','scenes','rules']`、`gd-runtime` `drawEntity` 原仅 circle/fill/rect、**无 sprite**。**【2026-06-19:Plan B U1(`db14845d`)已修——assets 改走 `sourceProject` 顶层抽取(非 strip)、`drawEntity` 加 `shape:'sprite'`→`drawImage` 受控面渲染;故运行时已能渲资产,唯九门仍不验资产对位。】**
|
||||
- **🔴 发现⑤:checkpoint 序列化潜伏**(评审 P0-5):图用 `SpringAIJacksonStateSerializer`(2 参 `StateGraph` ctor);`MysqlSaver` 序列化整 OverAllState 含历史键;e2e `saaCheckpointEnabled=false`+dataSource=null(~211/229)→ **量测从不触发序列化**;staging 默认 `saaCheckpointEnabled=true`(~190)→ 首 checkpoint 写/续跑才序列化新历史键。
|
||||
- **flip 当前无 env 钮**:`application-staging.yaml`(~173-185)只暴露 `enabled/api-key/worker-url/callback-secret`,**无 dispatcher/saa-source-mode/saa-model-protocol/llm-base**;`dispatcher` 默认 `http`(~160)、`saaSourceMode` 默认 `factory`(~198)、`llmBase` 默认 `http://100.64.0.8:3000`(~68);`AigcGenerateExecutor`(~151-152)`useSaa = "saa".equals(dispatcher) && saaDispatcher!=null`。`AigcExecutorProperties` 为 `@ConfigurationProperties("aigc.executor")`+`@Data`(props 已绑定,flip 是验证非新代码)。
|
||||
|
||||
---
|
||||
|
||||
## 可行性结论(外部 + 代码核证 · U-A/B/C 去险)
|
||||
|
||||
**FEASIBLE**:
|
||||
|
||||
- **依赖**:加 **bare `org.springframework.ai:spring-ai-anthropic`(无 version,`spring-ai-bom:1.1.2` 托管 @1.1.2)**,**非** `-starter`(避 `AnthropicAutoConfiguration` 抢注 bean)。**零 bean 冲突**(模块内无 `@Bean ChatModel`,模型全手装)。`mvn dependency:tree` 须证**无** `spring-ai-autoconfigure-model-anthropic` 传递入 + 无 `agentscope-core` + 无 spring-ai 漂离 1.1.2。
|
||||
- **协议**:`AnthropicChatModel` → `baseUrl=http://100.64.0.8:3000/anthropic`(`completionsPath` 默认 `/v1/messages`,**不复用 `stripV1Suffix`**);`AnthropicChatOptions.thinking(ENABLED, budget<maxTokens).maxTokens(≤128K)`;**阻塞 `.call()`**(流式 thinking bug #4407)。
|
||||
- **⚠️ 响应形(P0-4)**:每 content block 一 `AssistantMessage`(多 Generation);thinking 开时 **block0=thinking**。**必遍历 `resp.getResults()`**、答案取 text-role Generation、历史存有序 `List<AssistantMessage>`(thinking 先序)。`getMetadata()` 的 `thinking`/`signature`(皆 String)回读。
|
||||
- **序列化(P0-5)**:`AssistantMessage` 经 `SpringAIJacksonStateSerializer` 可往返,但 metadata 写为无类型 JSON、读回为 Map(thinking/signature 皆 String 可活);**`List<AssistantMessage>` 容器须经 `GenericListDeserializer` 保元素类型**(否则 `List<LinkedHashMap>` → 回放给 `AnthropicChatModel` 炸)。Anthropic 产 base `AssistantMessage`(非子类,通用 handler 适配)——须 checkpoint-ON 测试坐实。
|
||||
- **工具循环**:`DefaultToolCallingManager` + `while(hasToolCalls()) executeToolCalls(...)` 进单 `NodeAction`(无 agent-framework)。
|
||||
- **必修三险**:① 认证头方言(`.apiKey()` 发 `x-api-key`;网关若要 `Bearer` 则 401)= **spike 第一步**;② 流式 thinking bug #4407 → 只 `.call()`;③ 历史保真 → 新 client、**禁复用 `ExecutorLlmClient.normalizeJsonContent` 串扁平**。
|
||||
- **版本注**:本地 m2 仅有 anthropic 1.1.5(多 Generation 形从 1.1.5 反编译);1.1.2 首构网络拉取,**spike 须在实解 1.1.2 jar 上证多 Generation 形**(`resp.getResults().size()≥2`)。
|
||||
|
||||
---
|
||||
|
||||
## 关键技术决策(KTD)
|
||||
|
||||
1. **新 client + flag 旁挂,旧 OpenAI 路零动**:Anthropic 路新建,经 `aigc.executor.saa-model-protocol`(默认 `openai`)切;默认值字节零变、可瞬时回退。不改非 thinking 角色旧路。
|
||||
2. **历史保真是地基(多 Generation + 有界 + 可序列化)**:
|
||||
- **多 Generation 重组**:遍历 `resp.getResults()` 取有序块;答案取 text-role;历史项 = 每轮有序 `List<AssistantMessage>`(thinking 先序,符 Anthropic"assistant 轮以 thinking 起再 tool_use"铁律)。
|
||||
- **存储策略**:`OverAllState` 加历史键,**用 `ReplaceStrategy` + 节点内 read-append-write**(对齐 `appendAttempt`,不破 `ALL_KEYS`→ReplaceStrategy 不变量;P1-4)。
|
||||
- **有界 thinking-aware 截断(P1-3)**:保最近 K 轮完整 thinking、更早轮**降级保留 text+verdict 摘要、丢其 thinking**(守 token/成本,且保"thinking 先序"不变量);解 preserve-vs-bound 张力,非无界 append。
|
||||
- **可序列化(P0-5)**:历史键须经 `SpringAIJacksonStateSerializer` 往返保 `List<AssistantMessage>` 类型 + thinking/signature;checkpoint-ON 测试守。
|
||||
3. **并发 = 有界资源池、K 为正确性参数(非提速旋钮)**:
|
||||
- **池非锁单改(P1-5)**:`newSingleThreadExecutor`→有界池(size K) **且** `Semaphore(1)`→`Semaphore(K)` 守端口池;二者须同改(单线程下 Semaphore(K) 无效)。
|
||||
- **per-job 端口(P1-5)**:acquire 后 → `inputs.put("port",base+2i)`/`("cdpPort",base+i)` 替 `buildInputs` 常量;**`finally` 随 `playLock.release()` 释放**防超时/异常泄漏(180s 超时频发)。
|
||||
- **共享面审计 > 端口(P0-3)**:开 K>1 前审 `serve-and-play.sh` 的服根/`_shared/` 共读/`/tmp/*-$PORT.log`/`lsof|kill` 净场;端口池**整 job 生命期持槽**(净场不跨 job)。`serve-and-play.sh` 复用不改(harness 禁重写铁律);其每调 `mktemp -d` Chrome profile = 安全(已证)。
|
||||
- **K=正确性 + 双校准门**:保每条真玩实时(否则 `C_frame`/`E_live` 掉帧假阴);① 同 brief 集 K=1 vs K=N verdict 对等(时序);② **同款重复 ×N 全过一致**(结构竞态,P0-3)。**K=1 为记录 R1 基线权威,直至结构隔离证毕**。lili-mac(M1 Pro 8 核/32G、负载~6.6)K≈3-4。
|
||||
4. **量测路口径锁死 gameDefinition + 机制门警示**:e2e 暴露 `-Dsaa.e2e.sourceMode=gamedef`,R1 一律 gamedef。**但 ≥80% 仅证机制(九门对资产盲,P0-2)**,非产品质量;资产/好玩归 core-done R9/R11。
|
||||
5. **measure-first 闸控 U-C + 并行去险(P1-1)**:先 U-A/U-B + 8/8 driver → 跑有界并发基线 + Workflow-Opus 分析 → 收敛;**U-C(模型驱动工具循环)仅基线优化后仍 <80% 才全建**(守"自治在门内、裁决在门外")。**但与 U5 并行先跑 U-C 廉价 spike**(throwaway 探"M3 能否在本任务 写源→build→跑门→读 verdict→改 一轮工具循环",解 R-6 未证),把晚长杆从"未知可行"降为"已知可行待接线"。触发线预定:**≥2 轮单变量 A/B(同组 n≥10)仍 <80% 且根因指向"需模型自纠"**才全建 U-C。
|
||||
6. **Workflow-Opus = 收敛引擎、跑 6c6g**:全量 evidence(九门 verdict + 保留 thinking + 源 + 截图 + giveup-dump)经 Workflow fan-out(每失败款一 Opus agent 根因 → 综合排序优化)。Workflow 须 opt-in(创始人本轮指令即 opt-in),执行期触发。
|
||||
7. **flip 两步 + 协议同切 + 回退硬前置(P1-6/P1-7)**:① staging yaml 暴露 `dispatcher`/`saa-source-mode`/**`saa-model-protocol`**(`-D` 注入、默认不变);② `.env` 设 `saa`+`gamedef`+`anthropic` 切默认。**R1 证据记在用协议(量测==发布)**。**T1 回退验过 = step-2 硬前置**:失败(须改 gd-runtime/build-from-source)→ step-2 阻塞、列跨 Plan 关键依赖(非可延 TODO)。回退 = flag 切回 `http`(iife 主链 M2 80% 已证)。
|
||||
8. **运行时版本钉 + Plan B 冻结协调(P0-1 最高价值修)**:gen-done 量测执行 `gd-runtime.js`+`build-from-source.mjs`。**每轮基线/A-B 前钉二者 content-hash(git SHA),收敛窗口内 Plan B 对其变更 = 量测失效/重置基线事件**;与 Plan B 约定收敛窗口"不合 gd-runtime/build-from-source"冻结。一钉同时中和 P0-1 混淆 / P0-3 幻象基线 / P1-2 基底漂移。
|
||||
- **当前基底(2026-06-19)**:Plan B **U1 已推**(`db14845d` on `feat/mac-planB-coredone`,含引擎资产渲染 + asset-orphan 闭合 + 已过 ce-code-review 整改),**尚未合 `dev/2.0.0`**(仍 `0518d245`)。U1 = 我收敛的运行时基底。
|
||||
- **协调动作(推荐,由 6c6g/Plan B owner 驱动,我不擅合 mainline)**:U1 = 已完成的运行时单元 + 双线共享基底 → **先合 `dev/2.0.0` 并在我收敛窗口冻结**,我再 rebase 钉之 = 干净基底;优于我 rebase 到 Plan B 在飞分支(耦合其分支漂移)。**规划期我不需运行时在树**(量测在 ce-work/U5),故现暂不 rebase,待 U1 落 `dev/2.0.0`。
|
||||
|
||||
---
|
||||
|
||||
## 高层技术设计(HTD)
|
||||
|
||||
### 单元依赖 DAG + 收敛环
|
||||
|
||||
```mermaid
|
||||
flowchart TB
|
||||
subgraph FOUND["地基(用对 M3)"]
|
||||
U1["U1 Anthropic 协议 M3 client(U-A)<br/>flag 默认 openai · 多 Generation 重组"]
|
||||
U2["U2 thinking 历史保留(有界+可序列化) + 连续救场(U-B)"]
|
||||
U1 --> U2
|
||||
end
|
||||
subgraph HARNESS["量测地基"]
|
||||
U3["U3 gamedef 量测 + 有界并发池 harness<br/>sourceMode=gamedef · pool(K) · 双校准 · 共享面审计"]
|
||||
U4["U4 8/8 driver 决策树 + 上下文 + gatespec 约定"]
|
||||
end
|
||||
subgraph LOOP["收敛环(measure-first · 运行时钉)"]
|
||||
U5["U5 量测→Workflow-Opus 根因→优化→重测 至 ≥80%<br/>+ U-C 廉价 spike 并行去险 + 留出集 + 基线定标复核"]
|
||||
end
|
||||
U6["U6(闸控)模型驱动工具循环(U-C)<br/>仅 U5 ≥2 轮 A/B 仍<80% 才全建"]
|
||||
U7["U7 两步 flip + 协议同切 + 回退硬前置(003-U2)"]
|
||||
U2 --> U5
|
||||
U3 --> U5
|
||||
U4 --> U5
|
||||
U5 -. spike .-> U6
|
||||
U5 -- "<80%(满足触发线)" --> U6
|
||||
U6 --> U5
|
||||
U5 -- "≥80%(留出集)" --> U7
|
||||
U7 -. 产出消费 .-> PB["Plan B core-done 门"]
|
||||
PBfreeze["Plan B: gd-runtime/build-from-source<br/>收敛窗口冻结(KTD8)"] -. content-hash 钉 .-> U5
|
||||
```
|
||||
|
||||
### 收敛环时序(U5)
|
||||
|
||||
```mermaid
|
||||
sequenceDiagram
|
||||
participant Pin as 运行时钉(KTD8)
|
||||
participant Run as 有界并发跑(K 池)
|
||||
participant Ev as evidence 语料
|
||||
participant WF as Workflow(Opus) 6c6g
|
||||
participant Dev as 优化落地(U2/U4 + 调参)
|
||||
Pin->>Run: 钉 gd-runtime+build-from-source content-hash(变更则重置基线)
|
||||
Run->>Ev: n≥30 sourceMode=gamedef(mini-desktop 权威 K=1 记数 / Mac 快走查 K≈4 筛)
|
||||
Ev->>WF: 每失败款一 Opus agent 根因(门/thinking/源/截图)
|
||||
WF->>Dev: 综合聚类 → 排序优化项
|
||||
Dev->>Run: 改单变量 → 重测(同组 n≥10 A/B)
|
||||
Note over Run,Dev: 循环至 ≥80%(留出集);≥2 轮仍<80%→U6 工具循环 / 退路
|
||||
```
|
||||
|
||||
### flip + 回退状态(U7)
|
||||
|
||||
```mermaid
|
||||
stateDiagram-v2
|
||||
[*] --> http默认: 现状(dispatcher=http,iife,openai)
|
||||
http默认 --> 隔离验: 步① yaml 暴露 env + -D 注入 :48090(saa+gamedef+anthropic, checkpoint 开)
|
||||
隔离验 --> T1回退验: 隔离全过(health/端点/真玩/asset 烟测)+ smoke 绿
|
||||
T1回退验 --> 切默认: T1 过(存量 engineBundle 切 http 仍可玩)
|
||||
T1回退验 --> 阻塞: T1 失败(须改 Plan B 域)→ 跨 Plan 关键依赖
|
||||
隔离验 --> http默认: 隔离 fail(零改 live)
|
||||
切默认 --> [*]: 步② .env=saa+gamedef+anthropic 重部署 + smoke
|
||||
切默认 --> http默认: smoke 红/失败率超阈 → flag 切回(回退预案)
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 实现单元
|
||||
|
||||
> 骨架级:定 目标/范围/文件/验收。质量域 file 级逐字由 owner 在 U2/U4/U5 续做。测试机:lili-mac(快走查)+ mini-desktop(权威门)。
|
||||
|
||||
### U1. Anthropic 协议 M3 client(003-U1 U-A)
|
||||
|
||||
- **Goal**:接入"经 Anthropic `/v1/messages` 调 M3、thinking 开+干净分离、多 Generation 正确重组"的新 ChatModel,flag 旁挂(默认旧 OpenAI 路、字节零变)。
|
||||
- **Requirements**:R1 地基;003-U1 U-A。
|
||||
- **Dependencies**:无。
|
||||
- **Files**:`pom.xml`(加 bare `spring-ai-anthropic`);`.../saa/SaaStudioGraph.java`(新 `anthropicModel(...)` 工厂 + Anthropic base-url 推导,不复用 `stripV1Suffix`);`.../saa/SaaGraphDispatcher.java`(`ensureGraph` api 构造 ~316-317 旁挂 Anthropic 分支按 flag 选);`.../service/executor/AigcExecutorProperties.java`(加 `saaModelProtocol`(默认 `openai`)/`saaAnthropicBase`(默认 `http://100.64.0.8:3000/anthropic`)/thinking budget)。
|
||||
- **Approach**:`AnthropicApi.builder().baseUrl(anthropicBase).apiKey(NEWAPI_KEY)`;`AnthropicChatOptions.thinking(ENABLED, budget).maxTokens(≤128K).model("MiniMax-M3")`;**client 侧 fail-fast 断言 `budget<maxTokens`**(P2-3:勿信网关宽松、否则重演原截断 bug);阻塞 `.call()`。**spike 第一步 = 一次真 `.call()` 打网关**:401 则 `customHeaders(Authorization: Bearer …)`;并 `assert resp.getResults().size()>=2`(坐实 1.1.2 多 Generation 形,P2-2)。flag=openai 时不构造。
|
||||
- **Patterns**:现 `SaaStudioGraph.model()` 工厂 + dispatcher api-build 缝;NEWAPI_KEY 取 `docs/内网凭据与端点.md`。
|
||||
- **Test scenarios**:
|
||||
- happy(多 Generation):flag=anthropic 单条 `.call()` 返 `getResults().size()≥2`;text-role Generation 含答案、thinking Generation 含 `getMetadata().thinking` 非空。
|
||||
- 认证头:`x-api-key` 通过;否则 customHeaders Bearer 通过(记实际方言)。
|
||||
- fail-fast:`budget≥maxTokens` 时 client 构造即抛(不下发网关)。
|
||||
- 回归:flag=openai(默认)`mvn -pl game-module-aigc-server -am -DskipTests compile` 绿 + OpenAI 路字节等价;`mvn dependency:tree` 无 `spring-ai-autoconfigure-model-anthropic`/无 `agentscope-core`/无 spring-ai 漂离 1.1.2。
|
||||
- **Verification**:依赖树无漂移;默认 flag 零回归;anthropic flag 单调返 ≥2 Generation 含 thinking(日志 + gated 集成测试)。
|
||||
|
||||
### U2. thinking 历史保留(多 Gen + 有界 + 可序列化)+ 连续救场(003-U1 U-B)
|
||||
|
||||
- **Goal**:跨轮保留**完整有序模型响应**入历史(有界、可 checkpoint 序列化),救场改连续对话增量修。
|
||||
- **Requirements**:R1;003-U1 U-B;根因 v3 §3①。
|
||||
- **Dependencies**:U1。
|
||||
- **Files**:`.../saa/SaaStudioNodes.java`(`callAndRecord` ~235-248 / `judgeVision` ~673-707:**遍历 `getResults()` 重组**、存有序 `List<AssistantMessage>`、答案取 text-role;`generateNode`/`repairNode` 读写历史);`.../saa/SaaPrompts.java`(messages 改连续对话形;救场 prompt 去"重出完整模块"、改"针对反馈增量修正");`.../saa/SaaStudioGraph.java`(历史键登记 + **节点内 read-append-write**,不进 `ALL_KEYS` 全量 ReplaceStrategy 误判)。
|
||||
- **Approach**:`OverAllState` 加 `K_MSG_HISTORY`(List of 每轮有序消息组,thinking 先序);每轮 = 截断后历史 + 新 user(门 verdict/反馈);救场不重置历史、续上一轮 assistant 完整响应 + 失败 verdict。**有界**:保最近 K 轮完整 thinking、更早降级(text+verdict 摘要、丢 thinking),守 token/成本 + thinking 先序不变量。**gamedef 路与 iife 路各自历史键**。禁复用 `ExecutorLlmClient` 串扁平。
|
||||
- **Patterns**:`appendAttempt`(~298 节点内 append 范式);`SpringAIJacksonStateSerializer` 的 `AssistantMessageHandler`。
|
||||
- **Test scenarios**:
|
||||
- happy:连续两轮救场,第二轮含第一轮 assistant 完整响应(thinking 块未改/未丢)。
|
||||
- **深度(P1-3)**:max-repair 深度(5+3=8 轮)至少一次,thinking 先序不变量保持、API 不拒。
|
||||
- **序列化(P0-5)**:构 `MysqlSaver`(或 `MemorySaver` 经 `dataToBytes`/`dataFromBytes`)跑一 thinking 轮 → 强制序列化→反序列化 → 断历史仍 `List<AssistantMessage>`、thinking+signature 完整、可再下发。
|
||||
- **并发隔离(P0-3 联)**:两并发 run(一 gamedef 一 iife)历史键互不串。
|
||||
- 增量修:仅 H_progress 未过的源 + verdict,救场不打翻已过部分(diff 局部)。
|
||||
- token 追踪:U5 evidence 记**累计** token(含复发 thinking)。
|
||||
- 回归:iife 路(flag=openai)历史行为不变。
|
||||
- **Verification**:checkpoint-ON 序列化往返测试绿 + ≥2 轮且深度 8 thinking 保真 + 并发隔离 + 增量性人工抽检。
|
||||
- **Execution note**:先写"序列化往返保真"+"历史保真 ≥2 轮"失败测试再实现(最易静默回归 + 最易在生产才炸)。
|
||||
|
||||
### U3. gameDefinition 量测 + 有界并发池 harness
|
||||
|
||||
- **Goal**:让 e2e 门量对路(gamedef)且**有界并发可量**(K 为正确性参数),把 n≥30 从串行 ~4-8h 压到 ~1-1.5h,且不污染门、无结构竞态。
|
||||
- **Requirements**:R1(可客观运行硬门);发现②③。
|
||||
- **Dependencies**:无(与 U1/U2 并行可起)。
|
||||
- **Files**:`.../saa/SaaGraphDispatcher.java`(`graphExecutor` 改有界池(K) **且** `Semaphore(1)→Semaphore(K)` + 端口池;per-job `port`/`cdpPort` 注入替 `buildInputs` ~399-400 常量;`finally` 释放槽);`.../service/executor/AigcExecutorProperties.java`(加 `saaConcurrency` 默认 1 / `saaSourceMode` 暴露);`.../src/test/java/.../saa/SaaFullGraphE2eTest.java`(加 `-Dsaa.e2e.sourceMode`(默认 factory,量基线设 gamedef)/`-Dsaa.e2e.concurrency`(默认 1)/briefFile≥30);测试资源 `.../resources/saa-e2e-briefs-gamedef.txt`(≥30 条多品类 + **冻结钉 content-hash**)+ `saa-e2e-briefs-holdout.txt`(留出集,不进 A/B)。
|
||||
- **Approach**:端口池 = 基址 + 槽位(`4320+2i`/`9222+i`),`Semaphore(K)` 守池、acquire 拿空闲槽 → play 用该槽端口、**整 job 生命期持槽**(净场不跨 job)。`saaConcurrency` 默认 1 = 生产字节等价。**共享面审计**:开 K>1 前确认 `serve-and-play.sh` 服根/`_shared/` 共读只读安全、`/tmp/*-$PORT.log` 不撞、`lsof|kill` 不跨 job(`serve-and-play.sh` 复用不改)。**双校准门**:① 同 brief 集 K=1 vs K=N verdict 对等;② 同款重复 ×N 全过一致(结构竞态)。
|
||||
- **Patterns**:`SaaGenNodes` ~454-456 已读 state 端口;lili-mac `gamedef_quickcheck` per-game 端口错开(`QC_CONCURRENCY` 默认 4)实证。
|
||||
- **Test scenarios**:
|
||||
- happy:`-Dsaa.e2e.concurrency=4 -Dsaa.e2e.sourceMode=gamedef` 4 并发各异端口各产 verdict、无撞。
|
||||
- 校准①:同 6 款 K=1 vs K=4 verdict 对等(尤 `C_frame`/`E_live`)。
|
||||
- 校准②(结构隔离):同款重复 ×4 全过一致(无间歇竞态)。
|
||||
- 量对路:`sourceMode=gamedef` 真走 build-from-source(gamedef)非 iife。
|
||||
- 槽释放:play 超时(180s)/异常路端口槽 `finally` 释放、不泄漏。
|
||||
- 回归:默认(concurrency=1/sourceMode 未设)字节等价;`SaaFullGraphE2eGateLogicTest` 12/12。
|
||||
- **Verification**:端口分配/校准纯逻辑抽静态方法 + 模型无关单测;lili-mac 快走查证无撞 + 双校准过;**K=1 为记录基线权威直至结构隔离证毕**。
|
||||
- **Execution note**:default=1 先证字节零回归再开 K>1。
|
||||
|
||||
### U4. 8/8 driver 决策树 + 上下文 + gatespec 约定(003-U1 ③+②)
|
||||
|
||||
- **Goal**:补齐 driver 词汇下半 + 失败品类范例 + 门约定修对,512K 窗喂全。
|
||||
- **Requirements**:R1;根因 v3 §3②③。
|
||||
- **Dependencies**:无。
|
||||
- **Files**:`.../saa/SaaPrompts.java`(`DESIGN_SYSTEM` ~240-265 + `GAMEDEF_SYSTEM` ~147-235)。
|
||||
- **Approach**:现 5 driver(后二由在飞 `804e7617` 补)→ **补 `tap-pairs`/`key-cycle`/`drag-aiming`/`aim-fire` 至 8/8**;每 driver worked example(内联金样 `gd-archetypes.test.mjs:34-85`)+ 反例;**命门**:障碍跟踪类须显式产 `nextGap`/`nextPlatform` 等命名实体逐帧更新。**门约定(不改门本体)**:H/F 消除类用 `tap-pairs`+`score`;I_control 按 control-scheme 分型;`remaining` 语义经 `DESIGN_SYSTEM` java 范例(**不碰 gd-runtime**)。
|
||||
- **Patterns**:在飞 `804e7617`;`play.cdp.cjs` 8 driver(~207-214);金样 `match3` tap-pairs+score 过 H。
|
||||
- **Test scenarios**:编译绿(长串无破坏);离线 `extractGatespec` 对 8 driver 各样例解析正确;真选对率随 U5 量。
|
||||
- **Verification**:编译绿 + gatespec 解析单测;真效随 U5 闭环。
|
||||
|
||||
### U5. 收敛环:量测 → Workflow-Opus 根因 → 优化 → 重测 至 ≥80%
|
||||
|
||||
- **Goal**:有界并发量测 + 多 agent 根因,数据驱动把 gamedef 路过门率收敛到 ≥80%(留出集证)。**本里程碑收敛引擎。**
|
||||
- **Requirements**:R1;验收闭环 §6;根因 v3 §6。
|
||||
- **Dependencies**:U1+U2+U3(+U4)。
|
||||
- **Files**:Workflow 脚本(执行期落 session 目录,达标蒸馏 `.agents/skills/`);复用 U3 brief 文件 + evidence harness;无新生产代码。
|
||||
- **Approach**:
|
||||
1. **运行时钉(KTD8)**:每跑前钉 `gd-runtime.js`+`build-from-source.mjs` content-hash;收敛窗口 Plan B 变更 = 重置基线。
|
||||
2. **量基线**:`-Dsaa.e2e=1 -Dsaa.e2e.sourceMode=gamedef -Dsaa.e2e.n=30 -Dsaa.e2e.concurrency=K`——Mac 快走查(K≈3-4)筛、mini-desktop **K=1 权威记数**(标 minSuccessRate 旗 + 在用协议)。
|
||||
3. **基线定标复核(P1-2)**:首次干净 gamedef 基线后复核 80% 是否本路合适数(已 ≥85% → 门失信、转 rubric;~55% 有运行时天花板 → 早走退路)。
|
||||
4. **U-C 廉价 spike(P1-1)**:与本环并行跑 throwaway U-C 探针(M3 单轮工具循环可行性),去 R-6 险。
|
||||
5. **Workflow-Opus 分析**(6c6g,opt-in):每失败款一 Opus agent 根因 → 综合排序优化项。
|
||||
6. **优化落地**:改单变量、同组 n≥10 A/B。
|
||||
7. **重测 → 循环**至 ≥80%(**留出集 final 测,非调优集**);≥2 轮仍 <80% 且根因指向需自纠 → U6;天花板 → 退路。
|
||||
- **Patterns**:根因 v3「单变量 A/B 同组 n≥10」;效率策略 #8(Workflow 脚本);6c6g 角色。
|
||||
- **Test scenarios**:Test expectation: none —— 方法论/编排单元,其"验证"= R1 门(≥80% 实测,留出集)。
|
||||
- **Verification**:mini-desktop `SaaFullGraphE2eTest -Dsaa.e2e.minSuccessRate=0.8 -Dsaa.e2e.sourceMode=gamedef -Dsaa.e2e.n=30`(旗已设、协议记录、留出集)实测 ≥80%;运行时 hash 钉记录;各轮成功率/逐门/累计 token 留痕。
|
||||
|
||||
### U6.(闸控)模型驱动工具循环(003-U1 U-C)
|
||||
|
||||
- **Goal**:仅 U5 ≥2 轮 A/B 仍 <80% 且根因指向需自纠才全建——给 M3 工具(写源/build/跑九门/读 verdict/改),交错思维迭代。
|
||||
- **Requirements**:R1;003-U1 U-C;measure-first 闸控触发。
|
||||
- **Dependencies**:U5(触发线满足)+ U5 的 U-C spike(已证基本可行)。
|
||||
- **Files**:`.../saa/SaaStudioNodes.java`(新 react 式节点);`.../saa/SaaPrompts.java`(工具定义);`.../saa/SaaStudioGraph.java`(按 flag 接)。
|
||||
- **Approach**:`DefaultToolCallingManager` + `while(hasToolCalls()) executeToolCalls(...)` 进单 `NodeAction`;**九门作工具但 verdict 恒确定性**(模型不得自评过门 = "自治在门内、裁决在门外");`recursionLimit`+计数器封顶(控墙钟,每 run-gates 3.5-16min)。flag 旁挂、默认不启。
|
||||
- **Patterns**:Spring AI Tools「user-controlled execution」;`saa-graph-orchestration.md`「自治在门内裁决在门外」。
|
||||
- **Test scenarios**:happy(模型一轮 build→run-gates→读 verdict→改、结果回灌历史 thinking 顺序不破);确定性铁律(同款两次 run-gates verdict 一致;模型自称过但 verdict=fail 仍判 fail);封顶(达 recursionLimit 优雅 giveup)。
|
||||
- **Verification**:gated 集成测试证工具循环 + verdict 确定性;并入 U5 量 ≥80% 增益。
|
||||
|
||||
### U7. 两步 flip + 协议同切 + 回退硬前置(003-U2)
|
||||
|
||||
- **Goal**:质量达 ≥80% 后两步把默认从 `http`(iife,openai) 切 `saa`(gamedef,anthropic),含回退实证。
|
||||
- **Requirements**:R2;T1(回退硬前置);T2(两步分离)。
|
||||
- **Dependencies**:U5/U6 实测 ≥80%(留出集)。
|
||||
- **Files**:`game-cloud/huijing-server/src/main/resources/application-staging.yaml`(~173-185 加 `dispatcher: ${AIGC_EXECUTOR_DISPATCHER:http}` + `saa-source-mode: ${AIGC_SAA_SOURCE_MODE:factory}` + **`saa-model-protocol: ${AIGC_SAA_MODEL_PROTOCOL:openai}`**;**默认全不变**;`llm-base` **仅需 per-env 覆盖才加**否则去噪);`.../AigcExecutorProperties.java`(确认 env 绑定,@Data 已绑=验证非新代码)。
|
||||
- **Approach(两步,T2 分离)**:① **仅暴露 env** + `-D` 注入隔离实例 `:48090` 验 saa+gamedef+**anthropic**(**checkpoint 开**=首次真触序列化路 P0-5)+ **asset-orphan 最小烟测**(P0-2:至少一 sprite ref 可渲,即使全 R9 留 core-done),live `:48080` 零动(staging-ops §3);② **正式切默认**(staging `.env` 设三键,重部署 + `deploy/smoke-test.sh` 绿)。**R1 证据记在用协议(量==发)**。回退 = flag 切回 `http`。
|
||||
- **Patterns**:`staging-ops.md §3`(隔离 `-D` 注入、整机重部署安全变体);smoke 五链路。
|
||||
- **Test scenarios**:
|
||||
- happy:隔离实例 saa+gamedef+anthropic(checkpoint 开)一句话生成→入 feed→真玩 ≥1 款 CDP。
|
||||
- asset 烟测(P0-2):step-2 前至少一 sprite ref 在 play 截图可见(非"画面非空");不过 → step-2 不切。
|
||||
- 协议一致(P1-6):切后 `saa-model-protocol=anthropic` 生效(非 openai 默认);R1 证据协议字段=anthropic。
|
||||
- yaml 合并(P2-1):合并后 `aigc.executor` 块含两 Plan 全部期望键(config-presence 断言)、缩进合法。
|
||||
- 零回归:现 http 主链 smoke 五链路绿。
|
||||
- **T1 回退(硬前置)**:flip 后存量 SAA 产 engineBundle 切回 `http` 仍经 runtime/package 取包 + feed 真玩 ≥1 款 CDP。**失败(须改 gd-runtime/build-from-source)→ step-2 flip 阻塞、列跨 Plan 关键依赖(非 TODO)**。
|
||||
- T2 分离:步①只暴露 env、live 默认仍 http。
|
||||
- **Verification**:隔离全过(含 checkpoint 开 + asset 烟测)→ T1 过 → 切默认 → smoke 绿;DB/feed 行级核。
|
||||
|
||||
---
|
||||
|
||||
## 验收门(done 判据 · R1/R2)
|
||||
|
||||
- **R1(gen-done 核心 · 机制门口径)**:mini-desktop `SaaFullGraphE2eTest -Dsaa.e2e=1 -Dsaa.e2e.sourceMode=gamedef -Dsaa.e2e.n≥30 -Dsaa.e2e.minSuccessRate=0.8`(默认不带旗仅断 ≥1、不可作门;证据须打印 minSuccessRate 旗 + **在用协议=anthropic** + **运行时 content-hash** + **留出集标记**)实测 **≥80%**。**≥80% 仅证机制可玩(九门对资产/好玩盲,P0-2),非产品质量**——资产/玩法归 core-done R9/R11。与 smoke 互不替代。
|
||||
- **R2**:默认 `dispatcher=saa`+`saaSourceMode=gamedef`+**`saaModelProtocol=anthropic`**(量==发),一句话经 SAA 路生成→入 feed→真玩;worker 主链零回归;回退路验过(T1 为 step-2 硬前置)。
|
||||
- **gen-done = R1 ∧ R2**,独立里程碑、独立冻证。
|
||||
|
||||
---
|
||||
|
||||
## 范围边界
|
||||
|
||||
- **含**:003-U1(用对 M3 + 8/8 driver + 量测 harness + 收敛环)+ 003-U2(两步 flip + 协议同切 + 回退 + T1/T2)。
|
||||
- **不含(铁律)**:**编辑** `gd-runtime.js`/`build-from-source.mjs`/引擎资产渲染(= Plan B 003-U5);前端真验证/广度抽检/**资产渲染质量 rubric R9/R11**/M0 联合验收/最终 Phase-1 验收(= Plan B core-done)。**注**:gen-done 量测**执行**(非编辑)Plan B 两文件,故 ≥80% 是机制门、且需 KTD8 运行时钉。
|
||||
- **暂缓(次要 ④)**:best-of-N 生成、**生产侧多 job 并发**(生产 dispatcher 去串行是另一决策,带 split-brain/资源权重)——`saaConcurrency` 默认仍 1,本档只为**量测**开有界并发。
|
||||
- **退路(触发才走)**:升 `deepseek-v4-pro`/M3 强档 / Claude premium(质量优先记账);降九门判据/改判 Phase-1 门 = 末选创始人裁。
|
||||
- **worker `prompt.py` GAMEDEF_SYSTEM**:gen-done 不需(saa 路=量测/达标路、http=iife 回退路)→ 本档不补。
|
||||
|
||||
---
|
||||
|
||||
## 风险与依赖
|
||||
|
||||
- **R-1 认证头方言(最高首跑险)**:`.apiKey()` 发 `x-api-key`;网关若要 `Bearer` → 401。缓解 = U1 spike 第一步真 `.call()`,401 即 customHeaders Bearer。
|
||||
- **R-2 流式 thinking bug #4407 + thinking/tool_use 顺序**:只 `.call()`;多轮原样回放有序 `List<AssistantMessage>`;深度 8 测试。
|
||||
- **R-3 历史保真回归**:Anthropic 新 client 全程类型化消息;U2 先写保真失败测试。
|
||||
- **R-4 并发污染门 + 结构竞态(P0-3)**:K=正确性参数 + 双校准(时序对等 + 同款×N 一致)+ 共享面审计;超阈调小。
|
||||
- **R-5 算钟瓶颈**:K=4 下 n≥30 ~1-1.5h,A/B 多轮累积;mini-desktop 权威单实例串行排队。缓解 = Mac 快走查先筛、Workflow-Opus 减盲改轮。
|
||||
- **R-6 M3 工具-use 可靠性未实证**:U5 并行 U-C 廉价 spike 先证可行(P1-1)。
|
||||
- **R-7 ≥80% 可达性 + 标的未验(P1-2)**:gamedef 真基线 n=0(n=5 量错路);brief 集偏易自控(Goodhart)。缓解 = 基线定标复核 + brief 冻结钉 + 留出集 final 测 + 退路预置。
|
||||
- **R-8 Anthropic 多 Generation 形(P0-4)**:block0=thinking、`getResult()` 取错 → 抽取必败。缓解 = 遍历 `getResults()`、答案取 text-role、spike 证 `size()≥2`。
|
||||
- **R-9 checkpoint 序列化潜伏(P0-5)**:e2e harness checkpoint-off 隐藏序列化路 → staging(checkpoint 默认开)才炸。缓解 = U2 checkpoint-ON 往返测试 + U7 隔离验 checkpoint 开。
|
||||
- **R-10 运行时基底漂移(P0-1)**:gen-done 量测执行 Plan B 两文件,Plan B 收敛窗口改之 → A/B 混淆。缓解 = KTD8 content-hash 钉 + 收敛窗口冻结协调。
|
||||
- **跨 Plan 协调(与 Plan B 单向耦合)**:
|
||||
- **gd-runtime/build-from-source 冻结(P0-1/R-10)**:收敛窗口内 Plan B 不合并这两文件;每轮钉 hash。
|
||||
- `SaaFullGraphE2eTest`:本档**扩**(additive `-D` 旗/并发/sourceMode),Plan B **消费**(R4)→ 改动严格 additive。
|
||||
- `application-staging.yaml`:本档加 dispatcher/saa-source-mode/saa-model-protocol,Plan B U6 加 smoke 配置 → 同文件不同键,**合并须 config-presence 断言**(P2-1)、push 串行 `git ls-remote` 核。
|
||||
- **T1 回退依赖 Plan B 域(P1-7)**:T1 失败需改 gd-runtime/build-from-source → step-2 flip 阻塞、跨 Plan 关键依赖。
|
||||
- mini-desktop e2e **串行排队**:跑前确认无 Plan B 占用端口 4320/9222(flock 单实例)。
|
||||
- **依赖**:mini-desktop(权威九门 e2e,x86 prod-同构)+ lili-mac(快走查,M1 Pro 8 核/32G)+ 6c6g(Workflow 编排)+ `NEWAPI_KEY`(`docs/内网凭据与端点.md`)+ 已合 001 gamedef 运行时(`39c16630`)+ **收敛窗口 Plan B gd-runtime/build-from-source 冻结**。
|
||||
|
||||
---
|
||||
|
||||
## 需求追溯
|
||||
|
||||
| 目标 / origin | 单元 |
|
||||
|---|---|
|
||||
| 用对 M3 — Anthropic 协议 + thinking + 多 Generation 重组(U-A) | U1 |
|
||||
| 用对 M3 — 完整有界可序列化历史 + 连续救场(U-B / a3) | U2 |
|
||||
| gamedef 路量对 + 有界并发池可量(发现②③ / P0-3) | U3 |
|
||||
| 8/8 driver 决策树 + 上下文 + 门约定(③+②) | U4 |
|
||||
| 收敛到 ≥80%(R1,measure→Workflow-Opus→优化→重测 + 运行时钉 + 留出集) | U5 |
|
||||
| 模型驱动工具循环(U-C,闸控 + 并行 spike) | U6 |
|
||||
| 两步 flip + 协议同切 + 回退硬前置(R2 / T1/T2) | U7 |
|
||||
| 达标主杠杆=用对 agentic M3(主路)· 升强档/Claude premium(退路) | U1+U2+U5 + 退路 |
|
||||
|
||||
---
|
||||
|
||||
## 待确认(Open · 留待执行期/创始人)
|
||||
|
||||
- **U-C 触发线(P1-1,已预定)**:≥2 轮单变量 A/B(同组 n≥10)仍 <80% 且根因指向"需模型自纠"才全建 U6——确认该线。
|
||||
- **≥80% 是否本路合适数(P1-2)**:首次干净 gamedef 基线后复核(已 ≥85% 转 rubric / ~55% 走退路)。
|
||||
- **brief 分布对production input mix**:冻结 brief 集的品类分布与真实一句话输入分布的对齐度(留出集设计)。
|
||||
- **K 实测值**:lili-mac K≈3-4 / mini-desktop K 待双校准门定。
|
||||
- **认证头方言 / thinking budget**:U1 spike 第一步实测。
|
||||
- **step-2 flip 前 asset 烟测门槛(P0-2)**:最小"一 sprite 可渲"够否,还是需更强 asset 门(与 Plan B R11 对齐)。
|
||||
|
||||
---
|
||||
|
||||
## §6.8 双评审记录
|
||||
|
||||
本档已过 **Opus 对抗式 + Opus 可行性双评审**(代码核证,2026-06-19),整改 5 P0 + 7 P1 + 多 P2(见「双评审整改注记」,全 IN-DOCUMENT)。**Codex(codex-rescue) 11min 未出可用评审 → §6.8 Opus 单评回退**;ce-work 前可补跑 Codex 作 belt-and-suspenders。跨 Plan 协调项(gd-runtime/build-from-source 冻结、application-staging.yaml 合并、T1 跨 Plan 依赖、mini-desktop 串行排队)已列入「风险与依赖 · 跨 Plan 协调」,须与 Plan B owner 同步。
|
||||
Loading…
x
Reference in New Issue
Block a user