games-development-ai/deploy/infra/observability/otel-collector-config.yaml
lili 62ce5531b6
Some checks failed
contract-gates / contract-gates (push) Has been cancelled
docs-gate / docs-gate (push) Has been cancelled
feat(观测): 阶段四第一波观测栈 compose(五件+脱敏+MinIO后端,不起容器)
deploy/infra/observability/:OTel Collector+Prometheus+Tempo+Loki+Grafana docker-compose,照已签设计施工(opus 写 + fable 源码级终审)。
- 端口逐个不撞现有(Grafana 3002 避 new-api 3000/gitea 3001;Collector 4317-4318/Prom 9090/Tempo 3200/Loki 3100)
- 资源三上限(mem 合计 3.4G / cpus / cpu_shares=512 CPU 争用让路 new-api)
- 脱敏 redaction 三 pipeline 导出前(手机号/token/JWT/密钥)= 安全红线
- Tempo/Loki 存 MinIO 独立桶 obs-tempo/obs-loki(minio-init 幂等建、不碰 ragflow)
- 镜像全钉版本、密钥全走 .env 不入仓
终审验通:git 纯新增未 commit / compose config exit=0 / infra-minio 在 infra-shared 网 DNS 可解析(最大风险排除)/ grafana:11.4.0 经 daocloud mirror 拉取成功。起容器待创始人 go(碰生产 mini-infra)。

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-05 05:28:52 -07:00

149 lines
6.3 KiB
YAML

# =============================================================================
# OTel Collector 配置 —— 统一入口:收 OTLP → 脱敏 + 尾采样 + 批量 → 分发三库
# 设计 §4/§5 面三/§8:脱敏是进 Tempo/Loki 前的最后一道安全红线;尾采样错误+慢全采、
# 正常抽样;后端不可达时本地缓冲丢弃而非阻塞主链。
#
# 组件全部取自 opentelemetry-collector-contrib(redaction / tail_sampling 是 contrib 件,
# core 镜像没有,故 compose 用 -contrib 镜像)。
# =============================================================================
extensions:
# 健康探针:health-check.sh 探 :13133 判就绪。
health_check:
endpoint: 0.0.0.0:13133
receivers:
# OTLP 唯一入口。跨机信号源(game-cloud / 两条生成线,都在别处)经 Tailscale 推到
# mini-infra 的 4317(gRPC)/4318(HTTP)。监听 0.0.0.0 才能收到容器外/跨机流量。
otlp:
protocols:
grpc:
endpoint: 0.0.0.0:4317
http:
endpoint: 0.0.0.0:4318
processors:
# ── 内存护栏:软红线,超阈值即拒收/丢弃,防 Collector 自身 OOM 拖累同箱 new-api ──
# 达 limit 时开始拒绝新数据(back-pressure 上游 best-effort 上报会丢而非把 Collector 撑爆)。
memory_limiter:
check_interval: 1s
limit_percentage: 75
spike_limit_percentage: 15
# ── 尾部采样:一条 trace 的所有 span 齐了再决策(设计 §7)──────────────────
# 多策略取"或":错误全采 + 慢(>2s)全采 + 其余按 10% 概率采。既盯住所有异常/慢链路,
# 又把正常链路的存储量压到一成。decision_wait 是等齐 span 的窗口。
tail_sampling:
decision_wait: 10s
num_traces: 20000 # 决策期内最多缓存的 trace 数(内存换准确)
expected_new_traces_per_sec: 200
policies:
- name: errors-always # 错误链路:全采
type: status_code
status_code:
status_codes: [ERROR]
- name: slow-always # 慢链路(>2000ms):全采
type: latency
latency:
threshold_ms: 2000
- name: sample-normal # 其余正常链路:10% 抽样
type: probabilistic
probabilistic:
sampling_percentage: 10
# ── 脱敏:进库前最后一道安全红线(设计 §4/§8 第 7 条)────────────────────────
# 手机号 / token / 密钥 在写入 Tempo(span 属性)与 Loki(日志正文)前打码。
# allow_all_keys=true:不删属性(保留可观测性),只把命中键名或命中值的内容 mask。
# - blocked_key_patterns:键名像密钥的,整值打码(token/secret/password/api_key/…)。
# - blocked_values:值本身匹配敏感格式的打码(中国大陆手机号 / 常见密钥前缀 / JWT)。
# redact_all_types=true:非字符串(如 int 型手机号)也按字符串形态检查。
# summary=info:只记打码计数、不记键名,避免摘要本身泄信息。
redaction:
allow_all_keys: true
redact_all_types: true
blocked_key_patterns:
- "(?i).*(token|secret|password|passwd|api[_-]?key|access[_-]?key|authorization|cookie|credential).*"
blocked_values:
- "1[3-9]\\d{9}" # 中国大陆手机号
- "(?i)(sk|pk|api|key|token|bearer|ghp|xox)[-_][A-Za-z0-9]{16,}" # 常见密钥/令牌前缀
- "eyJ[A-Za-z0-9_-]{8,}\\.[A-Za-z0-9_-]{8,}\\.[A-Za-z0-9_-]{4,}" # JWT
summary: info
# ── 批量:降导出频次、提吞吐。放在 pipeline 末端、导出前。──
batch:
timeout: 5s
send_batch_size: 512
send_batch_max_size: 1024
exporters:
# ── 追踪 → Tempo(容器内 OTLP gRPC,同网明文 insecure)──────────────────────
# sending_queue = 本地内存缓冲;retry_on_failure = 后端抖动时重试。队列满即丢弃
# (不阻塞接收)——对应设计 §8"Collector 不可达/后端不可达时本地丢弃而非阻塞"。
otlp/tempo:
endpoint: tempo:4317
tls:
insecure: true
sending_queue:
enabled: true
queue_size: 5000
retry_on_failure:
enabled: true
initial_interval: 5s
max_interval: 30s
max_elapsed_time: 300s
# ── 指标 → Prometheus 出口(暴露 /metrics 在 8889,由 Prometheus 主动 scrape)──
# namespace=huijing:业务指标统一前缀 huijing_*(如 huijing_gen_task_total)。
# resource_to_telemetry_conversion:把 resource 属性(service.name 等)落成指标标签。
prometheus:
endpoint: 0.0.0.0:8889
namespace: huijing
resource_to_telemetry_conversion:
enabled: true
enable_open_metrics: true
# ── 日志 → Loki 的 OTLP 摄入端点(:3100/otlp,exporter 自动补 /v1/logs)──────
otlphttp/loki:
endpoint: http://loki:3100/otlp
sending_queue:
enabled: true
queue_size: 5000
retry_on_failure:
enabled: true
initial_interval: 5s
max_interval: 30s
max_elapsed_time: 300s
service:
extensions: [health_check]
telemetry:
# Collector 自身遥测暴露在 0.0.0.0:8888,供 Prometheus scrape(队列深度/丢弃数等,
# 用于验收第 6 条"agent 的 queue/drop 指标可见")。默认只在 localhost,须显式放开。
# 下面是 0.116 的 readers 写法;若该版本起容器时报 telemetry 配置错,退回老式(仍受支持):
# metrics:
# level: detailed
# address: 0.0.0.0:8888
metrics:
readers:
- pull:
exporter:
prometheus:
host: 0.0.0.0
port: 8888
pipelines:
# 追踪:限流 → 尾采样 → 脱敏 → 批量 → Tempo。脱敏在导出前、不可绕过。
traces:
receivers: [otlp]
processors: [memory_limiter, tail_sampling, redaction, batch]
exporters: [otlp/tempo]
# 指标:限流 → 脱敏(指标标签也可能带敏感值)→ 批量 → Prometheus 出口。
metrics:
receivers: [otlp]
processors: [memory_limiter, redaction, batch]
exporters: [prometheus]
# 日志:限流 → 脱敏(日志正文里有报错原文)→ 批量 → Loki。
logs:
receivers: [otlp]
processors: [memory_limiter, redaction, batch]
exporters: [otlphttp/loki]