fix(aigc): ExecutorLlmClient 显式 max_tokens=4096——对拍 llm_client.py 双侧固化,杜绝推理型通道吃光网关缺省额度致空content重试(C6.1实测22次空重试教训);雷一PromptResourceLoaderTest经核验本就版本无关设计,无需改
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
This commit is contained in:
parent
c2221bf2bc
commit
09d8678deb
@ -32,6 +32,14 @@ public class AigcExecutorProperties {
|
||||
public static final int LLM_MAX_RETRIES = 2;
|
||||
/** 单次 LLM HTTP 请求超时秒数 */
|
||||
public static final int LLM_REQUEST_TIMEOUT_SECONDS = 90;
|
||||
/**
|
||||
* 显式 max_tokens 上限(与编排器 llm_client.py:225 双侧对拍,取同款 4096)。
|
||||
* 为何显式固化(模型评估矩阵 2026-06-10 实证 + C6.1 实测):M2.7/M3/deepseek 系均为推理型通道,
|
||||
* 输出会先吃 reasoning_content;不显式传则依赖 new-api 网关缺省额度(隐式契约,且不同模型缺省不同),
|
||||
* 推理吃光额度后 message.content 为空 → 触发空 content 重试(C6.1 实测 22 次空重试;leg3 起显式给额 5 救 5)。
|
||||
* 4096 经 golden-regression v1.1.1/v1.1.2 抢救通道实证足够覆盖单款 GameConfig 输出,不取更大值以免无谓放大成本。
|
||||
*/
|
||||
public static final int LLM_MAX_TOKENS = 4096;
|
||||
|
||||
/**
|
||||
* 单次 LLM 调用链最坏耗时秒数 = 3 次尝试 ×90s + 退避(1s+2s) = 273s
|
||||
|
||||
@ -194,6 +194,9 @@ public class ExecutorLlmClient {
|
||||
body.put("model", properties.getLlmModel());
|
||||
body.put("temperature", AigcExecutorProperties.LLM_TEMPERATURE);
|
||||
body.putObject("response_format").put("type", "json_object");
|
||||
// 显式 max_tokens(对拍 llm_client.py:225):推理型通道不显式传会依赖网关缺省额度(隐式契约),
|
||||
// reasoning 吃光额度即得空 content 触发无谓重试——固化 4096 杜绝该截断/浪费类别(出处见 LLM_MAX_TOKENS 注释)
|
||||
body.put("max_tokens", AigcExecutorProperties.LLM_MAX_TOKENS);
|
||||
ArrayNode messages = body.putArray("messages");
|
||||
ObjectNode userMessage = messages.addObject();
|
||||
userMessage.put("role", "user");
|
||||
|
||||
Loading…
x
Reference in New Issue
Block a user