zizi d6e977a5ac feat(tier2): 图说对账补全核心引擎待补——n≥30 runbook基建+观测成本接线+L3软检(全加性/observe-only)
按 tier2 图说目标做缺口分析(8族逐元素比对),补齐 0号 spike 为过门收窄掉、
但图说明确要求的「核心引擎待补」项。全部加性/observe-only:金标冒烟仍 ACCEPT
(九门9/9+富游戏三门3/3,门一道没放松),真依赖下全链 import+自测+一款真 M3 跑验证通过。

G族(n≥30 runbook 执行基建):
- worker/config.py: build_model_openai 便宜档 client(deepseek 经 new-api OpenAI 兼容路,与 M3 Anthropic 路并存)
- worker/run_record.py: G4 采集字段表 → 可序列化 RunRecord(含退路树分流键 fail_system)
- worker/fallback_tree.py: 退路树五出口判定器(Q1–Q4 数字触发线,★阈值常量区待校准)
- batch_run.py / aggregate.py: model×variant×n 批跑(断点续跑/失败隔离)+ 矩阵聚合三图喂判定器

H族(观测/成本接线,把孤儿件缝进 run 主链):
- observability/newapi_pricing.py: 活读 new-api /api/pricing 倍率(取不到回落显式参数+告警)
- middleware.py: Tier2TraceMiddleware 挂 writer agent 最外层洋葱,ReAct 全事件旁路 ingest
- agent_loop/studio.py 收口: records→cost_for_run 折¥;真跑实测 cost_rmb=1.29(newapi-live)、trace 647事件 dropped=0
- contracts/trace/: additive trace 事件契约位(忠实 trace.py 落 sink 形状)

D族(L3 视觉软检接线,observe-only):
- agent_loop/studio.py: 收口调一次 M3 多模态(真截图+真玩取证→fun映射0-100),只写 verdict.L3,绝不参与 decision
- 真跑实测 L3 score=25 准确指出空心表现层;decision=fix 仍由 L1硬门/熔断裁、与 L3 无关(防 Goodhart 成立)

留后(不投机抢建):工作室 Agent Team/第二装载落库/控制面/Agent Service 等按 plan 决策②⑤ gate 到 B门后;
n≥30 等统计相是「跑」非「写」(批跑底座已就位);A-model 4插件复用待合并对账;4处图说 spec-drift 待 doc 线回写。
详见 tier2/HANDOFF.md「图说对账补全」节。

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-24 02:11:59 +00:00

160 lines
8.9 KiB
Python

"""observability/newapi_pricing.py —— tier2 富游戏自治线 · new-api 计费参数活读取(H3 成本台账接线)。
职责(对 docs/architecture/架构/生成引擎/tier2细节图说-H-观测与成本.md 图 H3「new-api quota 折¥」一步):
从 new-api 网关活读取折算所需的三件计费参数,喂给 observability.cost.cost_for_run 的
pricing / qpu / usd_rate 形参 —— 把「token → ¥」那一步从「显式硬传参」升级成「网关权威活值」。
权威成本源 = new-api(H3 钉死:成本不是估的,是从 new-api quota 口径折出来的),本模块就是接到这个口径的薄片。
① fetch_pricing() —— GET {base}/api/pricing → {model_name: {model_ratio, completion_ratio, cache_ratio}}。
这套倍率就是 new-api 计费引擎结算时用的同一套(与 wg1 cost.py / orchestrator newapi_cost.py 同口径,
已被 newapi-billing-plane-integration 验证);据它 + token 自算的 quota 与网关结算同公式(含缓存折扣)。
② fetch_status() —— GET {base}/api/status → (quota_per_unit, usd_exchange_rate);缺则取编译默认 500000 / 7.3。
③ fetch_pricing_params() —— 一步取齐三件,打成 cost_for_run 直接可用的 dict;**任何失败一律 best-effort 返回 None**
(不抛),让调用方回落到显式 pricing 参数并 log 警告——取价失败绝不中断生成主链(H1/H3 best-effort 铁律)。
为什么不直读 PG logs.quota(H3 图说提到的另一条路):
直读 new-api PostgreSQL 取每次调用的权威 quota(orchestrator newapi_cost.py 那条)要 mini-infra ssh + PG 口令 +
按时间窗关联,是「跑批事后对账」的重口径;tier2 这条线在 run 收口处要的是「当次 run 即时折一笔成本进 RunRecord」,
用 /api/pricing 倍率 + 当次 token 自算(与网关结算同公式)是最小且同口径的接法——pricing 倍率本身就是 quota 的来源。
跑批级的 PG 权威对账留给编排/分析侧(复用 orchestrator newapi_cost.py),本模块不重复造那条重链路。
tier2 .agent 红线(复用只走显式传参):
- base_url / api_key 经 worker.client(显式解析)传入,本模块不写死 base、不硬编码 key
(key 权威来源 = docs/内网凭据与端点.md NEWAPI_KEY;经 client.get_api_key 从 env/.env 读)。
- httpx 在函数内**惰性 import**:6c6g 静态校验环境无 httpx wheel,本模块仍能被 import(降级为「取价不可用、回落显式参数」),
真读只在装了依赖的 mini-desktop / Mac 上发生。
"""
from __future__ import annotations
from typing import Optional
# client 提供 get_api_key()(读 NEWAPI_KEY)与 resolve_base_url()(host 根 + 装代理旁路)。
# 包内/直跑兼容导入(直跑 observability 下脚本时 worker 包仍可解析)。
try:
from worker import client # type: ignore
except Exception: # pragma: no cover —— 直跑兜底:把 gen-worker/ 加进 sys.path 再取 worker.client
import sys
from pathlib import Path
sys.path.insert(0, str(Path(__file__).resolve().parents[1]))
from worker import client # type: ignore
# new-api /api/status 缺字段时的编译默认(与 wg1 cost.py / orchestrator newapi_cost.py 同源:
# 500000 quota = 1 USD;USD→¥ 汇率为标注假设、非 new-api 配置项,成本台账折算用)。
DEFAULT_QUOTA_PER_UNIT = 500000.0
DEFAULT_USD_RATE = 7.3
def fetch_pricing(base_url: str, api_key: str, *, timeout_s: float = 20.0) -> dict:
"""GET {base_url}/api/pricing → {model_name: {model_ratio, completion_ratio, cache_ratio}}。
这套倍率即 new-api 计费引擎结算用的同一套(cost.compute 据它 + token 按网关同公式自算 quota)。
:param base_url: new-api 网关 host 根(不带 /v1;凭据见 docs/内网凭据与端点.md NEWAPI_BASE_URL)。
:param api_key: NEWAPI_KEY(sk- token);经 Authorization: Bearer 头传。
:param timeout_s: 单次 HTTP 超时(秒)。
:return: {model_name: pricing_dict};网关无数据时返回 {}。
:raises: httpx / JSON 解析异常由调用方(fetch_pricing_params)best-effort 兜住,本函数不吞(便于上层记真因)。
"""
# 惰性 import:6c6g 无 httpx 时本模块仍可被 import(只是真读会在此抛 ImportError,由上层 best-effort 兜)。
import httpx # noqa: PLC0415 —— 故意函数内 import,避免模块级硬依赖 httpx
# base 归一去尾斜杠(防 'host:3000/' 拼成 'host:3000//api/pricing')。
base = (base_url or "").rstrip("/")
resp = httpx.get(
base + "/api/pricing",
headers={"Authorization": "Bearer " + api_key},
timeout=timeout_s,
)
resp.raise_for_status()
data = resp.json().get("data") or []
# new-api /api/pricing 返回 data 为模型数组,每项含 model_name + 各倍率;按 model_name 索引成 dict。
return {m["model_name"]: m for m in data if isinstance(m, dict) and m.get("model_name")}
def fetch_status(base_url: str, *, timeout_s: float = 20.0) -> tuple[float, float]:
"""GET {base_url}/api/status → (quota_per_unit, usd_exchange_rate);缺字段取编译默认。
/api/status 是公开端点(无需鉴权;wg1 cost.fetch_status 同口径);缺字段回落 500000 / 7.3。
:return: (quota_per_unit, usd_exchange_rate)。
:raises: httpx / JSON 异常由上层 best-effort 兜住,本函数不吞。
"""
import httpx # noqa: PLC0415 —— 函数内惰性 import(同 fetch_pricing 理由)
base = (base_url or "").rstrip("/")
resp = httpx.get(base + "/api/status", timeout=timeout_s)
resp.raise_for_status()
d = resp.json().get("data") or {}
qpu = d.get("quota_per_unit", DEFAULT_QUOTA_PER_UNIT)
usd_rate = d.get("usd_exchange_rate", DEFAULT_USD_RATE)
# 防脏值:非正数 qpu 会让 cost.compute 的 usd=quota/qpu 出问题,回落默认。
try:
qpu = float(qpu)
if qpu <= 0:
qpu = DEFAULT_QUOTA_PER_UNIT
except (TypeError, ValueError):
qpu = DEFAULT_QUOTA_PER_UNIT
try:
usd_rate = float(usd_rate)
except (TypeError, ValueError):
usd_rate = DEFAULT_USD_RATE
return qpu, usd_rate
def fetch_pricing_params(
*,
base_url: Optional[str] = None,
api_key: Optional[str] = None,
timeout_s: float = 20.0,
) -> Optional[dict]:
"""一步取齐折算三件套(pricing / qpu / usd_rate),打成 cost_for_run 直接可用的 dict。
**best-effort 铁律**:取价是成本台账的旁路,绝不能因网关抖动 / 6c6g 无 httpx / key 缺失而中断生成主链。
故任何异常一律捕获 → 返回 None + 落一条可追溯告警,让调用方回落到显式 pricing 参数(见 cost.cost_for_run)。
:param base_url: new-api host 根;None → client.resolve_base_url()(env NEWAPI_BASE_URL 或默认端点,已装代理旁路)。
:param api_key: NEWAPI_KEY;None → client.get_api_key()(从 env/.env 读;缺则抛,被本函数 best-effort 兜成 None)。
:param timeout_s: 单次 HTTP 超时。
:return: {"pricing": {...}, "qpu": float, "usd_rate": float} 或 None(取价失败,调用方回落显式参数)。
"""
try:
# base / key 解析:显式入参优先,否则经 client(红线:不在本模块写死 base、不硬编码 key)。
resolved_base = base_url or client.resolve_base_url()
resolved_key = api_key or client.get_api_key()
pricing = fetch_pricing(resolved_base, resolved_key, timeout_s=timeout_s)
qpu, usd_rate = fetch_status(resolved_base, timeout_s=timeout_s)
if not pricing:
# 取到空 pricing(网关无数据 / 模型表为空)也按取价失败处理,回落显式参数更安全。
print(
"[tier2-cost] new-api /api/pricing 返回空 pricing,"
"回落到调用方显式 pricing 参数。",
flush=True,
)
return None
return {"pricing": pricing, "qpu": qpu, "usd_rate": usd_rate}
except Exception as exc: # noqa: BLE001 —— best-effort:取价任何失败都不中断主链,只告警 + 回落
# 可追溯日志(错误路径铁律):记下取价失败真因,但不抛 —— 折算会回落到显式 pricing 参数。
print(
f"[tier2-cost] new-api 计费参数活读取失败(best-effort,回落显式 pricing 参数):"
f"{type(exc).__name__}: {exc}",
flush=True,
)
return None
if __name__ == "__main__": # pragma: no cover —— 本地自检:真连 new-api 取一次计费参数(需 httpx + 网关可达)
import json
params = fetch_pricing_params()
if params is None:
print("[newapi_pricing] 取价失败(见上方告警);真跑请在 mini-desktop / Mac 上装 httpx + 网关可达。")
else:
print(f"[newapi_pricing] qpu={params['qpu']} usd_rate={params['usd_rate']} "
f"models={len(params['pricing'])}")
# 抽样打印前 3 个模型的倍率(便于核对 MiniMax-M3 / deepseek 在不在表里)。
sample = dict(list(params["pricing"].items())[:3])
print(json.dumps(sample, ensure_ascii=False, indent=2))