一、技能重组(动作-对象命名) - 旧目录 clean/confirm/continuation/db/detect/embed/… 重组为 clean-book-text/decide-candidate/write-next-chapter/access-database/ check-content-consistency/embed-knowledge/…(git 识别为 rename,内容保持) - agents/*.md、AGENTS.md/CLAUDE.md 收编、example_skill 登记表同步新名 二、先审后入创作闭环(本次核心) 正文接受从"机械门一过就写正典"改为"机械门+语义审查双通过+用户批准+单事务原子提交", DB 级兜底,编排层跳步即被硬拒。 - candidate_cas.py + example_candidate_cas(109):持久化 CAS 状态链 - fact_delta.py + example_fact_delta/example_fact_ledger(106):结构化事实增量, 模型只提六型闭集增量+正文证据引文,仅用户批准的增量随正文同事务入账本 - projection_registry.py + example_projection_run(107):投影登记与恢复 - acceptance_state.py:接受前置实时状态重读 - lesson_registry.py + example_lesson(108):经验升格链,禁止自动升格 - DDL 105:example_candidate 增 semantic_status/semantic_report_sha256 - write_canonical.accept:语义兜底+同事务合并增量+登记投影; run_writer_pipeline/persist_writer_run/run_writer_semantic_detector/step2 接入全链 - claude_runtime:兼容新 CLI modelUsage 信息字段 三、审查修复(独立子代理四维审查后) - 事实增量 propose→approve 翻态正道,不撞唯一键 - 冻结配置探针重刷(CLI 2.1.211→2.1.231 漂移),profileSha256/adapterVersion 再登记 - 可视化合同悬空路径/五六空间矛盾、 SoT 旧技能名漂移、行尾空白清理 测试:离线 65 套 + 真实库集成 5 套(CAS/接受故障注入/事实增量/投影/经验升格)+ 回放 79 项全绿。 创作内容(docs/design、生成正文 artifacts)按"框架与创作分开"未入本提交。
107 lines
5.7 KiB
Python
107 lines
5.7 KiB
Python
#!/usr/bin/env python3
|
||
"""clean-book-text Skill:LLM 探测执行器——每窗一次受治理调用,产出 deletions JSON。
|
||
|
||
读 clean_prep 产出的 manifest.json,逐窗经 llm.chat_governed 探测垃圾段(只报逐字原文,不改写),
|
||
写 /tmp/muse-clean/<work>/deletions-NNN.json。断点续跑:已有产物的窗自动跳过。
|
||
模型降级与额度治理已上收 llm.chat_governed(全局 BUDGET_CHAIN + 5h 额度窗),本脚本不自写降级链。
|
||
删除动作不在本脚本:由 clean_apply 守卫裁决执行。
|
||
"""
|
||
import json
|
||
import pathlib
|
||
import sys
|
||
import time
|
||
|
||
import click
|
||
|
||
# 统一走 call-content-model 受治理入口(额度窗/全局降级链/熔断 + trust_env/重试/<think>剥离/JSON 容错都在那边)
|
||
sys.path.insert(0, str(pathlib.Path(__file__).resolve().parents[2] / "call-content-model" / "scripts"))
|
||
from llm import chat_governed, extract_json # noqa: E402
|
||
|
||
OUT = pathlib.Path("/tmp/muse-clean")
|
||
|
||
# 探测 prompt:与 SKILL.md 模板同源;要求纯 JSON 输出便于机器解析
|
||
PROMPT = """你是网文正文清洗探测器。下面是《{title}》第 {a}–{b} 章的原文(每章以【第N章 | 章题】标记行开头)。
|
||
|
||
找出**所有非正文垃圾**:
|
||
- 书站广告及变体(如「一秒记住♂粒÷小÷说→网」「天才一秒记住本站地址」)
|
||
- 网址/域名残留、"首发""手机版阅读""无弹窗"类导流语
|
||
- 求票拉票/求订阅收藏/打赏鸣谢段
|
||
- 章尾作者话(ps:… / PS:… / 附:…)、上架感言、请假条类整段
|
||
- 乱码水印、明显不属于故事的插入符号串
|
||
|
||
输出规则(只输出一个 JSON 对象,禁止任何其他文字):
|
||
{{"deletions": [{{"chapter_order": 章序号, "exact": "待删段逐字原文", "reason": "简短理由"}}]}}
|
||
|
||
硬性要求:
|
||
- `exact` 必须与原文**逐字一致**(含标点、空格),单项 ≤300 字;同一段垃圾一项,不合并多段;
|
||
- 只报确定是垃圾的,拿不准的不报;**情节正文一个字都不许报**;
|
||
- 【第N章 | 章题】标记行不报(即使含求票字样);
|
||
- 若该窗没有垃圾,输出 {{"deletions": []}}。
|
||
|
||
原文开始:
|
||
{text}"""
|
||
|
||
|
||
@click.command()
|
||
@click.option("--work-id", type=int, required=True)
|
||
@click.option("--win", "wins", type=int, multiple=True, help="指定窗号(可多次);不给则全部窗")
|
||
@click.option("--model", default="MiniMax-M3", show_default=True,
|
||
help="首选入口模型(兼容 hint);实际用哪个模型与降级由 llm.chat_governed 全局额度策略决定")
|
||
@click.option("--force", is_flag=True, help="已有产物也重跑")
|
||
def main(work_id, wins, model, force):
|
||
d = OUT / str(work_id)
|
||
manifest = json.loads((d / "manifest.json").read_text())
|
||
title = manifest["title"]
|
||
targets = [m for m in manifest["windows"] if not wins or m["win"] in wins]
|
||
total_in = total_out = n_del = n_skip = 0
|
||
for m in targets:
|
||
f_out = d / f"deletions-{m['win']:03d}.json"
|
||
if f_out.exists() and not force:
|
||
click.echo(f"win-{m['win']:03d} 已有产物,跳过(--force 重跑)", err=True)
|
||
continue
|
||
text = pathlib.Path(m["file"]).read_text()
|
||
t0 = time.time()
|
||
prompt = PROMPT.format(title=title, a=m["from"], b=m["to"], text=text)
|
||
# 降级与额度治理已上收 llm.chat_governed(全局 BUDGET_CHAIN + 5h 额度窗):撞内容安全/
|
||
# 模型不可用由它沿全局链自动换模型并按窗预算/调用数治理;本脚本不自写降级链。
|
||
# 单窗失败不中断整书——全链耗尽或输出无法解析时失败关闭:该窗不清洗、不返回假成功。
|
||
content, usage, used_model = chat_governed(prompt, model=model, caller="clean")
|
||
if used_model is None:
|
||
# 治理链全部耗尽(多为上游敏感词拦截):写空产物占位(含跳过原因),
|
||
# apply 端窗产物齐备可继续,审计可追——绝不拿空内容当成功
|
||
f_out.write_text(json.dumps(
|
||
{"deletions": [], "skipped": "治理链全部耗尽(多为上游敏感词拦截)"},
|
||
ensure_ascii=False, indent=1))
|
||
n_skip += 1
|
||
click.echo(f"win-{m['win']:03d}(第{m['from']}–{m['to']}章)探测跳过(治理链耗尽),该窗不清洗")
|
||
continue
|
||
try:
|
||
data = extract_json(content)
|
||
except ValueError as e:
|
||
# 成功调用但输出无法解析为 JSON:失败关闭,该窗不清洗、不返回假成功
|
||
f_out.write_text(json.dumps(
|
||
{"deletions": [], "skipped": f"LLM 输出无法解析为 JSON({type(e).__name__})"},
|
||
ensure_ascii=False, indent=1))
|
||
n_skip += 1
|
||
click.echo(f"win-{m['win']:03d}(第{m['from']}–{m['to']}章)输出解析失败,该窗不清洗")
|
||
continue
|
||
dels = data.get("deletions", data if isinstance(data, list) else [])
|
||
total_in += usage.get("prompt_tokens", 0)
|
||
total_out += usage.get("completion_tokens", 0)
|
||
f_out.write_text(json.dumps({"deletions": dels, "model": used_model},
|
||
ensure_ascii=False, indent=1))
|
||
n_del += len(dels)
|
||
tag = "" if used_model == model else f"(治理降级至 {used_model})"
|
||
click.echo(f"win-{m['win']:03d}(第{m['from']}–{m['to']}章)检出 {len(dels)} 段{tag} "
|
||
f"→ {f_out.name}({time.time() - t0:.0f}s)")
|
||
click.echo(f"探测完成:{len(targets)} 窗共检出 {n_del} 段(跳过 {n_skip} 窗);"
|
||
f"token in={total_in:,} out={total_out:,}")
|
||
|
||
|
||
if __name__ == "__main__":
|
||
try:
|
||
main()
|
||
except RuntimeError as e:
|
||
click.echo(f"[llm错误] {e}", err=True)
|
||
sys.exit(1)
|