修复升格停书bug——出场章int/str混排炸+并集虚增立卡计数(深空窗113实证,四处归一化+6用例)

This commit is contained in:
zizi 2026-07-18 20:30:48 +08:00
parent 26f5dad589
commit 8e9e97c913
2 changed files with 42 additions and 6 deletions

View File

@ -358,6 +358,23 @@ def _chapter_sort_key(ch):
return CHAP_BIG
def _int_chaps(chaps):
"""出场章归一化为整数集合(窗113 实证 bug 修复):模型偶尔把章号输出成字符串(如 "508"),
与库内 int 章号一起 sorted 会炸('<' not supported between int and str,深空窗113 当场停书);
且 set 并集把 "508" 和 508 当两个值、会虚增跨章计数误判立卡门槛。这里把数字串强制转 int、
非数字(脏值/区间)丢弃——出场章只应是单章整数(区间只出现在里程碑「章」,不在出场章)。
bool 是 int 子类,显式排除以防 True/False 混入被当章号。"""
out = set()
for c in (chaps or []):
if isinstance(c, bool):
continue
if isinstance(c, int):
out.add(c)
elif isinstance(c, str) and c.strip().isdigit():
out.add(int(c.strip()))
return out
def _infer_lifecycle(text):
"""从台阶文本启发式推断生命周期枚举(模型未给或非法「周期」时的降级填充)。
诚实边界:这是关键词启发式、非精确判定;新抽取由提示词强制模型直接给枚举,此路仅兜底。"""
@ -639,16 +656,17 @@ def new_card(conn, work_id, win_no, ent, milestone_types=None):
out.append(f"[窗{win_no}] {core}")
seen.add(core)
fields0[k] = out
chaps_int = sorted(_int_chaps(ent.get("出场章", []))) # 归一化 int(窗113 修复):登场兜底与落库共用
# 洞② 登场兜底:该型合同含「演变历程」但抽取结果无登场里程碑时,机械补一条登场(出场章 min + 摘要)——
# 模型倾向只填当前态、漏建登场,此为「提示词硬约束 + 机械兜底」双保险里的机械那一保险。
if milestone_types and ent.get("型") in milestone_types:
fields0["演变历程"] = _debut_milestone(
fields0.get("演变历程") or [], ent.get("一句话摘要", ""), ent.get("出场章", []), win_no)
fields0.get("演变历程") or [], ent.get("一句话摘要", ""), chaps_int, win_no)
payload = {"type": ent["型"], "名称": raw,
"别名": [x for x in ([_clean_alias(a) for a in ent.get("别名", [])] + extra_alias) if x],
"一句话摘要": ent.get("一句话摘要", ""),
"字段": fields0,
"出场章": sorted(set(ent.get("出场章", []))),
"出场章": chaps_int,
"来源": f"升格@窗{win_no}", "状态": "草稿",
"目标库": "本书作品库", "可见范围": "本书私有",
"_work_id": work_id}
@ -763,7 +781,7 @@ def _classify_new_name(ent, name_map, presence):
and nm != ex and (nm in ex or ex in nm)), None)
if sub_hit:
return "substr", sub_hit
chaps = set(ent.get("出场章", []))
chaps = _int_chaps(ent.get("出场章", [])) # 归一化 int:避免 "508"/508 混算虚增跨章计数
hist = presence.get((t, nm), set())
if len(chaps | hist) >= 2: # 跨章(含跨窗合计)→ 立卡
return "new", None
@ -1109,7 +1127,7 @@ def run(work_id, max_windows, max_calls, model, redo_window, semantic_on):
(work_id, key, nm, win_no, TENANT))
continue
# kind in ("new","presence"):跨章立卡 / 单章龙套留档
chaps = set(ent.get("出场章", []))
chaps = _int_chaps(ent.get("出场章", [])) # 归一化 int(窗113 修复)
hist = presence.get((ent.get("型", ""), nm), set())
if kind == "new": # 跨章(含跨窗合计)→ 立卡
if hist: # 用留档补足初卡出场章
@ -1282,7 +1300,8 @@ def run(work_id, max_windows, max_calls, model, redo_window, semantic_on):
continue
p = conn.execute("SELECT draft_payload FROM muse_knowledge_draft WHERE id=%s",
(did,)).fetchone()[0]
p["出场章"] = sorted(set(p.get("出场章", [])) | chs)
# 归一化 int(窗113 修复):库内旧 payload 可能残留字符串章号,与 chs(已 int)混排会炸
p["出场章"] = sorted(_int_chaps(p.get("出场章", [])) | chs)
conn.execute("UPDATE muse_knowledge_draft SET draft_payload=%s WHERE id=%s",
(json.dumps(p, ensure_ascii=False), did))
conn.execute(

View File

@ -161,6 +161,22 @@ def _fake_contracts():
return c
# ── ⑦ 出场章归一化 _int_chaps(窗113 实证 bug:int/str 混排炸 + 并集虚增计数)──
def test_int_chaps():
# 核心复现:模型给字符串章号 "508" 与库内 int 507 混合,旧代码 sorted 直接炸
mixed = [507, "508", 509]
got = pu._int_chaps(mixed)
check("intchaps-混类型归一为int", got == {507, 508, 509})
check("intchaps-归一后可安全排序", sorted(got) == [507, 508, 509])
# 并集不再虚增:{"508"} 与 {508} 旧代码算 2 个(误判跨章立卡),归一后算 1 个
union = pu._int_chaps(["508"]) | {508}
check("intchaps-并集不虚增计数", union == {508} and len(union) == 1)
# 脏值/区间/bool 丢弃(出场章只应是单章整数)
check("intchaps-丢非数字与区间", pu._int_chaps(["420-423", "abc", None, 12]) == {12})
check("intchaps-排除bool", pu._int_chaps([True, False, 5]) == {5})
check("intchaps-空输入", pu._int_chaps([]) == set() and pu._int_chaps(None) == set())
# ── ⑥ 三个提示词含关键纪律语句(洞②登场必须 / 洞③≤40字 + 体系级 / 洞②归并去重)──
def test_prompts_disciplines():
c = _fake_contracts()
@ -190,6 +206,7 @@ def test_prompts_disciplines():
if __name__ == "__main__":
for fn in (test_classify_new_name, test_debut_milestone, test_clean_milestone_guard,
test_merge_material, test_build_embed_text_type_fix, test_prompts_disciplines):
test_merge_material, test_build_embed_text_type_fix, test_int_chaps,
test_prompts_disciplines):
fn()
print(f"\n全部离线自测通过:{_passed} 项(未连库、未发任何网络/嵌入/LLM 调用)")