Add cross-model veto of hair-gap background residue (default on)

Colour-drifted background trapped between hair strands defeats every
single-signal defence: the chroma key reads it as foreground (bgc ~0.07),
the segmenter backs it, ViTMatte rates it opaque, and post-hoc removal is
a proven dead end (it shreds the hair volume the same pixels belong to).
A second, trimap-free matting model (BiRefNet_HR-matting) is the only
tested model that separates this residue from the subject, so its opinion
is fused in as a veto: min-fusion that may only LOWER alpha, restricted to
the background-hued bright suspect zone (proj >= 3, L >= 45, feathered)
and gated by primary-alpha confidence (0.70 -> 0.95 ramp) so soft wisps
and dark hair are exempt by construction.

- settings/config/CLI: cross_check block, --cross-check/--no-cross-check
- alpha_post.cross_check_alpha after clean_alpha; second opinion reuses
  BiRefNetSegmenter; saved to debug as cross_check_alpha.png
- chroma.bg_hue_projection extracted and shared with the trimap
- docs: methodology.md (new), hair_gap_artifacts.md (investigation log)

Verified: cross-check ON reproduces the visually-reviewed B1gate
prototype byte-for-byte on TestImage3; --no-cross-check reproduces the
previous baseline byte-for-byte; pink-bg FixImage1 face untouched.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
This commit is contained in:
2026-07-04 19:52:24 +08:00
parent 639fdf01b6
commit adbbb9d620
11 changed files with 474 additions and 15 deletions
+10
View File
@@ -9,6 +9,7 @@ RGB input
-> chroma bg-confidence (keyed to the auto-detected or given colour)
-> [optional] semantic segmentation mask
-> trimap -> ViTMatte -> alpha cleanup
-> cross-model veto (second matting opinion on bg-hued residue)
-> pymatting foreground -> despill -> RGBA PNG -> QA previews
```
@@ -28,6 +29,15 @@ Both share pymatting foreground estimation and a colour de-spill. There is **no*
green-contamination rescue / recolour layer — with clean source images it is
unnecessary, so it was removed.
Both also run a **cross-model veto** by default: a second, trimap-free matting
model (`ZhengPeng7/BiRefNet_HR-matting`) may only *lower* alpha, only on
background-hued bright pixels the primary result is confident about — this clears
colour-drifted background residue trapped between hair strands that the chroma
key, the segmenter and ViTMatte all read as foreground. Costs one extra model
download (~0.9 GB) and one inference pass per image; disable with
`--no-cross-check` (see the `cross_check` config section, and
`docs/hair_gap_artifacts.md` for the analysis behind it).
The segmentation trimap defaults to `directional` mode (chroma + seg + a hue-direction
split: it keeps a background-coloured garment such as a white shirt while dropping a
background-hued residual such as blue trapped between hair strands). Switch with
+42 -1
View File
@@ -3,7 +3,7 @@ from __future__ import annotations
import numpy as np
from .deps import require_cv2
from .settings import AlphaPostSettings
from .settings import AlphaPostSettings, CrossCheckSettings
def enforce_trimap(alpha: np.ndarray, trimap: np.ndarray) -> np.ndarray:
@@ -56,6 +56,47 @@ def suppress_alpha_by_chroma(
return np.clip(out, 0.0, 1.0)
def cross_check_alpha(
alpha: np.ndarray,
second_alpha: np.ndarray,
proj: np.ndarray,
lightness: np.ndarray,
trimap: np.ndarray,
settings: CrossCheckSettings,
) -> np.ndarray:
"""Veto falsely-confident alpha with an independent second matting opinion.
Background residue trapped between hair strands drifts in colour until the
chroma key reads it as foreground, the segmenter backs it, and the primary
matte rates it opaque -- every single-signal defence fails. A second model
that does separate it (see docs/hair_gap_artifacts.md) may pull alpha DOWN
(min-fusion; it can never raise it), and only where
- the pixel is background-hued and not dark (the suspect zone), and
- the primary alpha is high (``gate_lo -> gate_hi`` ramp): pixels the
pipeline already renders soft (outer wisps) are exempt by construction.
Inside the zone this deliberately overrides the trimap-FG clamp -- the
residue it exists to clear is mostly trimap-FG. Sure background cannot be
disturbed: min-fusion keeps alpha 0 at 0.
"""
if not settings.enabled:
return alpha
cv2 = require_cv2()
zone = (
(proj >= settings.proj_min)
& (lightness >= settings.l_min)
& (trimap != 0)
)
weight = zone.astype(np.float32)
if settings.feather_sigma > 0:
blur = cv2.GaussianBlur(weight, (0, 0), settings.feather_sigma)
weight = np.where(zone, 1.0, np.clip(blur, 0.0, 1.0)).astype(np.float32)
gate = weight * _smoothstep(alpha, settings.gate_lo, settings.gate_hi)
out = alpha * (1.0 - gate) + np.minimum(alpha, second_alpha) * gate
return np.clip(out, 0.0, 1.0).astype(np.float32)
def _remove_small_components(mask: np.ndarray, min_area: int) -> np.ndarray:
cv2 = require_cv2()
count, labels, stats, _ = cv2.connectedComponentsWithStats(mask.astype(np.uint8), 8)
+15
View File
@@ -39,6 +39,21 @@ def convert_color_spaces(rgb: np.ndarray) -> tuple[np.ndarray, np.ndarray]:
return rgb_f, lab
def bg_hue_projection(
lab: np.ndarray, lab_center: tuple[float, float, float]
) -> np.ndarray:
"""Per-pixel projection of Lab chroma (a*, b*) onto the background hue direction.
Positive = displaced toward the background hue, ~0 for neutral pixels. Returns
zeros when the background is near-neutral (no meaningful hue direction).
"""
center = np.asarray(lab_center, dtype=np.float32)
mag = float(np.hypot(center[1], center[2]))
if mag < 1e-3:
return np.zeros(lab.shape[:2], dtype=np.float32)
return (lab[..., 1] * center[1] + lab[..., 2] * center[2]) / mag
def parse_hex_color(text: str) -> tuple[float, float, float]:
"""Parse ``#RRGGBB`` (or ``RRGGBB``) into an RGB triple normalised to 0..1."""
s = text.strip().lstrip("#")
+2
View File
@@ -29,6 +29,7 @@ def main(
unknown_radius_ratio: float | None = typer.Option(None, "--unknown-radius-ratio", min=0.0),
fg_safe_radius_ratio: float | None = typer.Option(None, "--fg-safe-radius-ratio", min=0.0),
despill: bool | None = typer.Option(None, "--despill/--no-despill"),
cross_check: bool | None = typer.Option(None, "--cross-check/--no-cross-check", help="Second-opinion veto of background-hued residue between hair strands (default: on; costs one extra model inference)"),
trimap_mode: str | None = typer.Option(None, "--trimap-mode", help="Trimap mode (segmentation pipeline): directional | seg | directional-hard-bg"),
seg_backend: str | None = typer.Option(None, "--seg-backend", help="Segmentation backend: birefnet (default) | anime-seg"),
) -> None:
@@ -45,6 +46,7 @@ def main(
unknown_radius_ratio=unknown_radius_ratio,
fg_safe_radius_ratio=fg_safe_radius_ratio,
despill=despill,
cross_check=cross_check,
trimap_mode=trimap_mode,
seg_backend=seg_backend,
)
+7 -1
View File
@@ -7,6 +7,7 @@ from typing import Any, TypeVar
from .settings import (
AlphaPostSettings,
ChromaSettings,
CrossCheckSettings,
DespillSettings,
ForegroundSettings,
ModelSettings,
@@ -48,7 +49,7 @@ def _update_dataclass(instance: T, values: dict[str, Any] | None) -> T:
def settings_from_dict(data: dict[str, Any]) -> PipelineSettings:
data = dict(data)
screen_color = data.pop("screen_color", None)
allowed_sections = {"chroma", "trimap", "alpha_post", "foreground", "despill", "model", "segmentation"}
allowed_sections = {"chroma", "trimap", "alpha_post", "cross_check", "foreground", "despill", "model", "segmentation"}
unknown_sections = sorted(set(data) - allowed_sections)
if unknown_sections:
raise ValueError(f"Unknown config section(s): {', '.join(unknown_sections)}")
@@ -57,6 +58,7 @@ def settings_from_dict(data: dict[str, Any]) -> PipelineSettings:
chroma=_update_dataclass(ChromaSettings(), data.get("chroma")),
trimap=_update_dataclass(TrimapSettings(), data.get("trimap")),
alpha_post=_update_dataclass(AlphaPostSettings(), data.get("alpha_post")),
cross_check=_update_dataclass(CrossCheckSettings(), data.get("cross_check")),
foreground=_update_dataclass(ForegroundSettings(), data.get("foreground")),
despill=_update_dataclass(DespillSettings(), data.get("despill")),
model=_update_dataclass(ModelSettings(), data.get("model")),
@@ -100,6 +102,9 @@ def override_settings(settings: PipelineSettings, **overrides: Any) -> PipelineS
despill_updates: dict[str, Any] = {}
if overrides.get("despill") is not None:
despill_updates["enabled"] = overrides["despill"]
cross_check_updates: dict[str, Any] = {}
if overrides.get("cross_check") is not None:
cross_check_updates["enabled"] = overrides["cross_check"]
seg_updates: dict[str, Any] = {}
if overrides.get("device") is not None:
seg_updates["device"] = overrides["device"]
@@ -121,6 +126,7 @@ def override_settings(settings: PipelineSettings, **overrides: Any) -> PipelineS
chroma=chroma,
trimap=_update_dataclass(trimap, trimap_updates),
alpha_post=alpha_post,
cross_check=_update_dataclass(settings.cross_check, cross_check_updates),
foreground=foreground,
despill=_update_dataclass(despill, despill_updates),
model=_update_dataclass(model, model_updates),
+42 -5
View File
@@ -6,13 +6,23 @@ from pathlib import Path
import numpy as np
from .alpha_post import clean_alpha, enforce_trimap, suppress_alpha_by_chroma
from .chroma import compute_bg_confidence, parse_hex_color
from .alpha_post import (
clean_alpha,
cross_check_alpha,
enforce_trimap,
suppress_alpha_by_chroma,
)
from .chroma import (
bg_hue_projection,
compute_bg_confidence,
convert_color_spaces,
parse_hex_color,
)
from .despill import despill
from .foreground import estimate_foreground_rgb
from .io import load_rgb, save_gray, save_rgb, save_rgba, write_text
from .qa import BACKGROUND_COLORS, composite, make_qa_grid, save_previews
from .settings import PipelineSettings
from .settings import PipelineSettings, SegmentationSettings
from .trimap import (
fuse_trimap,
fuse_trimap_directional,
@@ -27,6 +37,7 @@ class MattingPipeline:
self.settings = settings
self._runner: ViTMatteRunner | None = None
self._segmenter = None
self._cross_checker = None
def _segment(self, rgb: np.ndarray) -> np.ndarray:
if self._segmenter is None:
@@ -35,6 +46,20 @@ class MattingPipeline:
self._segmenter = make_segmenter(self.settings.segmentation)
return self._segmenter.mask(rgb)
def _second_opinion(self, rgb: np.ndarray) -> np.ndarray:
if self._cross_checker is None:
from .segmentation import BiRefNetSegmenter
check = self.settings.cross_check
self._cross_checker = BiRefNetSegmenter(
SegmentationSettings(
model_name=check.model_name,
device=self.settings.segmentation.device,
input_size=check.input_size,
)
)
return self._cross_checker.mask(rgb)
def _predict_alpha(self, rgb: np.ndarray, trimap: np.ndarray, bg_confidence: np.ndarray) -> tuple[np.ndarray, str]:
if self.settings.model.matting_method == "chroma":
return trimap_to_alpha_seed(trimap, bg_confidence), "chroma"
@@ -89,8 +114,6 @@ def _run_image(
if mode == "seg":
trimap, trimap_stats = fuse_trimap(seg_mask, bg_confidence, settings.trimap)
elif mode in ("directional", "directional-hard-bg"):
from .chroma import convert_color_spaces
lab = convert_color_spaces(rgb)[1]
trimap, trimap_stats = fuse_trimap_directional(
seg_mask, bg_confidence, lab, model.lab_center, settings.trimap
@@ -113,6 +136,18 @@ def _run_image(
raw_alpha=raw_alpha if alpha_source == "vitmatte" else None,
)
alpha = clean_alpha(alpha, trimap, settings.alpha_post)
second_alpha = None
if settings.cross_check.enabled:
lab = convert_color_spaces(rgb)[1]
second_alpha = pipeline._second_opinion(rgb)
alpha = cross_check_alpha(
alpha,
second_alpha,
bg_hue_projection(lab, model.lab_center),
lab[..., 0],
trimap,
settings.cross_check,
)
foreground = estimate_foreground_rgb(
rgb, alpha, bg_confidence, model, settings.foreground
)
@@ -137,6 +172,8 @@ def _run_image(
save_gray(debug / "trimap.png", trimap)
if seg_mask is not None:
save_gray(debug / "seg_mask.png", seg_mask)
if second_alpha is not None:
save_gray(debug / "cross_check_alpha.png", second_alpha)
save_gray(debug / "alpha.png", alpha)
save_rgb(debug / "foreground_rgb.png", corrected_rgb)
save_gray(debug / "color_mask.png", color_mask)
+19
View File
@@ -63,6 +63,24 @@ class AlphaPostSettings:
suppress_raw_hi: float = 0.98
@dataclass(frozen=True)
class CrossCheckSettings:
# Cross-model veto: a trimap-free matting model gives a second opinion that
# may only LOWER alpha (min-fusion), only on background-hued, non-dark pixels
# (the suspect zone) where the primary result is confident (gate_lo->gate_hi
# alpha ramp). Clears colour-drifted background residue between hair strands
# that chroma, segmentation and the primary matte all read as foreground;
# already-soft wisps and dark hair are exempt by construction.
enabled: bool = True
model_name: str = "ZhengPeng7/BiRefNet_HR-matting"
input_size: int = 2048
proj_min: float = 3.0 # bg-hue projection above which a pixel is suspect
l_min: float = 45.0 # Lab lightness below which a pixel is exempt (dark hair)
feather_sigma: float = 2.0 # Gaussian feather of the zone boundary, in px
gate_lo: float = 0.70 # primary alpha below this -> fully exempt
gate_hi: float = 0.95 # primary alpha above this -> fully vetoable
@dataclass(frozen=True)
class ForegroundSettings:
enabled: bool = True
@@ -117,6 +135,7 @@ class PipelineSettings:
chroma: ChromaSettings = ChromaSettings()
trimap: TrimapSettings = TrimapSettings()
alpha_post: AlphaPostSettings = AlphaPostSettings()
cross_check: CrossCheckSettings = CrossCheckSettings()
foreground: ForegroundSettings = ForegroundSettings()
despill: DespillSettings = DespillSettings()
model: ModelSettings = ModelSettings()
+3 -8
View File
@@ -2,6 +2,7 @@ from __future__ import annotations
import numpy as np
from .chroma import bg_hue_projection
from .deps import require_cv2
from .settings import TrimapSettings
@@ -137,14 +138,8 @@ def fuse_trimap_directional(
chroma_unknown = ~chroma_bg & ~chroma_fg
core = seg_mask >= settings.seg_core_threshold
# Directional chroma: projection of (a*, b*) onto the background chroma
# direction. Positive = colour displaced toward the background hue.
lab_c = np.asarray(lab_center, dtype=np.float32)
mag = float(np.hypot(lab_c[1], lab_c[2]))
if mag < 1e-3:
proj = np.zeros(shape, dtype=np.float32)
else:
proj = (lab[..., 1] * lab_c[1] + lab[..., 2] * lab_c[2]) / mag
# Directional chroma: positive = colour displaced toward the background hue.
proj = bg_hue_projection(lab, lab_center)
bg_hued = proj >= settings.bg_hue_proj_min
bg = chroma_bg | ((seg_mask < settings.seg_low) & ~chroma_fg)
+16
View File
@@ -46,6 +46,22 @@ alpha_post:
suppress_raw_lo: 0.85
suppress_raw_hi: 0.98
cross_check:
# Cross-model veto: a second, trimap-free matting model may only LOWER alpha
# (min-fusion), only on background-hued bright pixels the primary result is
# confident about — clears colour-drifted background residue between hair
# strands that chroma, segmentation and ViTMatte all read as foreground.
# Costs one extra model (~0.9 GB download) and one inference pass per image.
# Disable with enabled: false or --no-cross-check.
enabled: true
model_name: ZhengPeng7/BiRefNet_HR-matting
input_size: 2048
proj_min: 3.0 # bg-hue projection above which a pixel is suspect
l_min: 45.0 # Lab lightness below which a pixel is exempt (dark hair)
feather_sigma: 2.0 # zone-boundary feather, px
gate_lo: 0.70 # primary alpha below this -> fully exempt
gate_hi: 0.95 # primary alpha above this -> fully vetoable
foreground:
enabled: true
method: ml
+128
View File
@@ -0,0 +1,128 @@
# 发丝间隙残留瑕疵:问题分析与候选优化方向
状态:**待验证**(2026-07 记录)。这是当前管线剩下最顽固的质量问题。
## 问题描述
发丝间隙和个别镂空洞中残留蓝白色瑕疵点。颜色已偏离背景色(如 #CFEFFF),
是背景与发丝的稀释混合色。在深色合成背景上最显眼。
## 为什么现有机制全部够不着(失败链,每环均已实测)
1. **chroma 抓不到**:混合色漂移后 bgc ≈ 0.04,远低于 sure_fg 阈值 0.12,
chroma 直接判为前景。
2. **seg 抓不到**:瑕疵被头发包在轮廓内部,seg 置信度极高
(seg_loose 拉到 0.95 仍保留,sweep 已验证)。
3. **ViTMatte 分不出来**:几像素宽的亮色细缝,与反光亮发丝/头发软体积在
局部纹理上几乎不可区分,模型给出中等偏高 alpha。
4. **chroma_suppress 救不了**:压制由 bgc 驱动(smoothstep 0.35→0.80),
bgc 0.04 不触发;matte-confidence gate(639fdf0)对它们无关——
还没到门就被 bgc 挡在外面。
5. **后处理摘除已证死路**:speck 检测器(L 5078 / proj≥3.5 / 暗邻域 /
小连通域)能命中,但同批像素同时是真实头发的软体积,压 alpha 或挖洞
都会打碎右腰/胳膊间的发簇。多背景评审判定 baseline 最好。
**核心教训:问题不在"检测不到",而在"动作是毁灭性的"。** 只要动作是
"把 alpha 压向 0",检测误报就直接毁头发,而误报无法消除——瑕疵与发丝
软体积在颜色、拓扑、alpha 邻域上全部重叠(拓扑检测器 v1–v4 全失败)。
## 已证伪、不要再试
- 任何"检测 + 压 alpha / 挖洞"的变体(见 5)。
- trimap 阈值/带宽类调参(seg_loose 0.100.95 两轮 sweep 无增益)。
- 拓扑/连通域/到背景距离类检测(口袋与内部拓扑不可分)。
- 换 seg 模型(BEN2 与 BiRefNet 打平,已废弃)。
## 候选方向(按优先级)
### A. 局部放大重推理 —— 已验证,证伪(2026-07,不要落地)
根因假设是:原分辨率下缝隙仅 2–3 像素,模型证据不足;放大重推理应能让
ViTMatte 把缝隙 alpha 判低。先导验证(TestImage3,4 个瑕疵密集 crop,
2×2 对照:放大 1x/2x/4x × trimap 原样/可疑 FG 降级 unknown/降级+膨胀)
结果:
- 只放大(trimap 原样):speck 中位 alpha 1.000 → 1.000,零变化(2x、4x 同)。
- 放大 + 降级:中位仍 0.98–0.99,与不放大只降级打平甚至更差。
- 放大 + 降级 + 膨胀 11px(给模型最大自由):3/4 crop 中位 0.880.96,
最好的 crop 也只到 0.57——仍远高于视觉消除所需;且 4x 常比 2x 更高
(更多证据让模型**更确信**是前景)。
- 视觉:右腰/胳膊间关键区域各变体与 baseline 几乎不可区分。
诊断补充:瑕疵点 46.6% 是 trimap-FG(rule 2:bgc 0.07 中位 ≤0.12 且
seg 0.94 中位),其余 53.4% 在 unknown 带但 ViTMatte raw 中位 0.989。
**结论:不是分辨率/证据不足,是画风层面歧义**——anime 发丝间的亮色细条
在 ViTMatte 眼里就是头发高光/软体积,给它更多像素只会更确信。任何以
ViTMatte 重判为动作的方案(含分块放大整图版)都不必再做。这同时大幅
降低方向 B 的期望:歧义在样式而非模型容量,换 matting 模型大概率同判
(参考 BEN2 教训);若试 B,先用同样的 4 crop 先导验证,不要直接接管线。
### B. 换更强的 matting 模型 —— 先导验证部分通过(2026-07,待视觉评审)
**同族换权重:证伪。** ViTMatte Distinctions-646(唯一零接入候选;AEMatter
不在 HF,FBA/DiffMatte 仅权重无代码)与 Com-1K 判决一致(demote 后
speck 中位 0.870.97),同架构换训练数据无效。
**跨族第二意见:通过。** BiRefNet_HR-matting(trimap-free,2048 输入,
trust_remote_code 加载方式与现有 seg 相同)是**第一个把瑕疵和主体拉开
的模型**:speck med 0.74 / p10 0.36,而暗发核心 0.996、衬衫 1.0、
软发丝 med 0.18(本就该半透明)。直接替换不可行(软发丝太薄、丢硬保证),
但可作**交叉模型融合**:
zone = (proj>=3.0) & (L>=45) & (trimap!=0) # 可疑色区,向外羽化
gate = smoothstep(alpha_base, 0.7, 0.95) # 只否决假自信像素
B1: alpha = lerp(alpha, min(alpha, hrmat), zone*gate)
B2: alpha = lerp(alpha, alpha*smoothstep(hrmat, 0.5, 0.95), zone*gate)
置信门是关键:第一版没有它,软发丝被全灭(med 0.477→0.000);加门后
speck mean 0.769→0.714(B1)/0.673(B2),发核心/衬衫/软发丝逐像素不变。
视觉:右腰发簇完好,缝隙蓝白明显消退。效果图 Outputs/dirB_hrmat/。
落地成本:+1 个模型(~0.9GB)+1 次推理(cuda ~2s)。遗留:两模型都判
~1.0 的瑕疵子集(p90=1.0)不受影响。
**视觉评审结论(2026-07-04):B1gate 胜出。** 用户多区域评审后选保守的
min 融合;B2 的意见锐化(remap 0.5→0.95)被否——数字更低(0.673 vs
0.714)但视觉代价不值。落地形态因此更简:无 remap 参数,第二意见原样
`min()`,可调参数只剩 zone(proj_min/L_min/羽化)与置信门(lo/hi)。
**已落地(2026-07-04,默认开启)。** 新增 `cross_check` 设置块 +
`--cross-check/--no-cross-check`;pipeline 在 clean_alpha 后插入
`cross_check_alpha`(alpha_post.py),第二意见加载复用 BiRefNetSegmenter。
验证:开启时 TestImage3 输出与评审原型逐字节一致;`--no-cross-check`
旧 baseline 逐字节一致;FixImage1(粉底)脸部无损、发缘粉残留被清,
唯一副作用是羽化权重外溢到暗豁免区边缘(弓弦零星像素最多 -0.5,中位
无变化,视觉连续)——若在意可把羽化后的权重对 L<l_min 强制清零。
**SEMat (SAM2) 已评估,不接入(2026-07-04)。** 在 WSL(Ubuntu-26.04,
conda BGfilter,torch 2.12.1+cu130,detectron2 源码装,权重 XiaRho/SEMat
970MB,1024 letterbox + bbox 提示,推理 ~1s)跑了同协议先导验证:
speck med 0.860(弱于 HR-matting 的 0.737),hair/shirt 1.0,wisp 0.250。
与 HR-matting 在瑕疵上相关系数 0.913——看到的是同一批;在 HR≥0.95 的
顽固子集(1723/4116 px)上 SEMat med 恰为 1.000,零补充。min() 组合仅
0.737→0.699,不值第二个模型的成本。至此瑕疵顽固子集已被**四个模型族**
(ViTMatte Com-1K/D646、BiRefNet_HR-matting、SEMat-SAM2)一致判为前景
——后处理侧到此为止,顽固子集只能走生成端(D 的色相轴)。
### C. 只动颜色,不动 alpha(A 证伪后,后处理侧唯一剩余选项)
瑕疵刺眼在于颜色(蓝白)而非覆盖率。把 despill 作用域扩到内部口袋,
沿背景色 Lab 方向压掉蓝味,alpha 不动——几何上不可能毁头发。
上限有限:蓝白变灰白,点还在,只是不扎眼。
注意:与已废弃的 recolor 层同属"后修颜色"气味,但那是重建断发的结构性
修补,这里只是现有 despill 的作用域调整,性质轻得多。
### D. 源头:背景色选择(生成端,非管线改动;A 证伪后优先级最高)
**饱和度轴已关闭(2026-07)**:提高饱和度会让生成模型把背景色画进角色
(头发染绿)——这正是当年 recolor/rescue 层存在的理由,修复染色源图已证
死路。饱和度实为一个把污染在生成侧(染色)和后处理侧(瑕疵点)之间搬运
的滑块,两端均已证不可行,中间只是按比例混合。
**剩余可动维度:色相轴(待验证)**。瑕疵不可分的根因是蓝白泄漏与 anime
发丝高光(冷蓝/白,画风常量)同色。保持雾蓝级低饱和、把色相换到画风
不用于高光的方向(如淡薄荷绿 #CFF5DF):染色风险与淡蓝同级,但泄漏
带绿味(proj 朝绿)、高光仍蓝白(proj≈0)——"不可分"判决只对蓝背景
成立,绿色相下 proj 驱动的压制可能直接可用。角色无绿色系内容,撞色
风险也低于蓝。验证:同构图生成淡绿背景版,查 (1) 染色是否同级,
(2) 泄漏与高光的 proj 分布是否拉开。
+190
View File
@@ -0,0 +1,190 @@
# BgFilter 方法论(Methodology)
> 面向平色背景 AI 角色图像的离线抠图算法。交叉模型否决(§7)经视觉评审
> 后已落地,默认开启。
## 1. 问题设定与总体框架
给定一幅在**单一平色背景**上生成的角色图像 $I \in \mathbb{R}^{H\times W\times 3}$,
目标是估计逐像素不透明度 $\alpha \in [0,1]^{H\times W}$ 与前景色
$F$,输出高质量 RGBA。任务难点集中在发丝、半透明边缘、衣物镂空,
以及与背景色相近的主体区域(白衬衫、浅肤色)。
算法的总体框架是**三方分权**:
- **色度先验(chroma)** 掌握"什么颜色是背景"——提供连续的背景置信度场,
并锐化平背景边界;
- **语义分割(segmentation)** 掌握"哪里是主体"——决定拓扑(把染色发丝留在
前景、把透视孔洞留在背景);
- **matting 模型(ViTMatte)** 掌握"边界处的混合比例"——只在双方都不确定的
未知带内做精细 alpha 推理。
三方通过一张三值 trimap($\{0, 128, 255\}$,即背景/未知/前景)交换决定权;
连续证据(置信度、模型原始输出)则绕过三值接口,在后处理阶段以软方式继续
参与(§4)。管线存在一个退化形态:关闭分割后成为纯 chroma + ViTMatte 的
两方管线,本文以三方管线为主线。
## 2. 背景颜色先验与置信度场
### 2.1 背景色自动探测(带失败闸门)
无人工输入时,从图像边框条带(宽度约 $4\%$ 边长)探测背景色:对边框像素的
Lab 值做粗直方图(bin 宽 6),取众数 bin 的均值为峰中心,再以半径 12 聚出
背景簇。设两道质量闸门:
1. 簇须占边框像素 $\geq 55\%$;
2. 四角(各取 $3\%$ 短边见方)中位色与簇中心距离 $\leq 20$ 的须 $\geq 3$ 角。
任一不满足即抛出异常而非硬算——渐变、双色、主体压角等场景应显式失败,
避免对错误的背景色抠图。探测器与后续算法解耦,仅输出一个颜色。
### 2.2 背景颜色模型
对样本集合(自动模式取边框簇像素;人工输入 hex 时取 Lab 距目标 $\leq 25$
的像素,优先边框,不足则回退全图最近 10%)做稳健统计:中位数为中心
$c_{\mathrm{lab}}, c_{\mathrm{rgb}}$,以绝对偏差 75 分位 $\times 1.4826$
为尺度 $\sigma$(设下限,防止过窄的方差把轻微压缩噪声推出背景)。
### 2.3 背景置信度场
对每个像素以双空间高斯核取较大者:
$$C = \max\!\Big(\exp\big(-\tfrac{1}{2}(d_{\mathrm{lab}}/\sigma_{\mathrm{lab}})^2\big),\;
0.85\,\exp\big(-\tfrac{1}{2}(d_{\mathrm{rgb}}/2.5)^2\big)\Big)$$
其中 $d_{\mathrm{rgb}}$ 为按通道 $\sigma$ 归一的欧氏距离。$C$ 是全管线的
色度主证据:参与 trimap、alpha 压制与前景估计。
## 3. 语义–色度三信号三值 Trimap
### 3.1 三个信号
- $C$:背景置信度(§2.3),阈值 $\tau_{bg}=0.92$、$\tau_{fg}=0.12$ 把像素分为
chroma-背景 / chroma-前景 / chroma-未知三档;
- $s$:分割软掩码(BiRefNet,1024 输入;可切换 anime-seg),
阈值 $\tau_{low}=0.15$、$\tau_{core}=0.60$;
- $\rho$:**色调方向投影**。设背景 Lab 色度向量为 $(a_c, b_c)$,则
$$\rho(p) = \frac{a_p a_c + b_p b_c}{\lVert (a_c, b_c) \rVert}$$
$\rho \geq 4.0$ 记为"偏背景色调"(bg-hued)。$\rho$ 与 $C$ 互补:$C$ 度量
到背景色的**距离**,$\rho$ 度量色偏的**方向**——白衬衫距背景不远但方向中性,
发间蓝残留距背景也不近但方向朝蓝,二者由 $\rho$ 分开。
### 3.2 融合规则(directional 模式,默认)
$$\mathrm{BG}: \quad C \geq \tau_{bg} \;\lor\; (s < \tau_{low} \land C > \tau_{fg})$$
$$\mathrm{FG}: \quad C \leq \tau_{fg} \land s \geq \tau_{low}$$
$$\text{Rule 3}: \quad \underbrace{\tau_{fg} < C < \tau_{bg}}_{\text{chroma-未知}} \land\; s \geq \tau_{core} \land \rho < 4.0 \;\Rightarrow\; \mathrm{FG}$$
其余像素保持未知。规则体现**证据分级原则**:规则 2 中分割仅需
$s \geq 0.15$,因为它有 chroma 的独立佐证;规则 3 中分割单独作证
(chroma 弃权),故门槛提高到 $0.60$,且被 $\rho$ 一票否决——偏背景色调的
高置信像素(发丝间的背景残留)**留在未知带**,交由 §4 的压制处理,而非
硬定前景。变体 `seg`(无色调切分)与 `directional-hard-bg`(bg-hued 直接
判背景,激进,可能误伤冷色阴影下的白布)按参数选用。
最后沿分割轮廓(loose 掩码 $s \geq 0.08$ 的膨胀−腐蚀差)保留一圈细未知带
(不覆盖 chroma-背景),供 matting 模型抗锯齿;带宽刻意小,以保住指缝等
小孔洞。
## 4. Trimap 引导的 Alpha 推理与后验修正
### 4.1 推理与硬钳制
ViTMatte(base,Composition-1K 权重)以 RGB + trimap 为输入输出原始
alpha $\alpha_{raw}$。随后 `enforce_trimap` 施加硬保证:
$$\alpha[\mathrm{BG}] \equiv 0, \qquad \alpha[\mathrm{FG}] \equiv 1$$
这使"确定背景恒透明"成为结构性质而非统计性质——模型幻觉无法漏入
已判定的背景。ViTMatte 不可用时可回退 chroma 种子
$\alpha = \mathrm{clip}(1 - C)$(按参数启用)。
### 4.2 带 matte 置信门的色度压制
未知带内残存的背景色(发丝缝隙、孔洞死角)由色度证据压除。记
$S(x; e_0, e_1)$ 为 smoothstep,压制系数为
$$\lambda = S(C;\, 0.35,\, 0.80)\; \cdot\; \big(1 - S(\alpha_{raw};\, 0.85,\, 0.98)\big)$$
$$\alpha' = \alpha \cdot (1 - k\,\lambda), \quad k=1, \quad \text{仅作用于未知带}$$
第二个因子是 **matte 置信门**:ViTMatte 自信不透明($\alpha_{raw} \to 1$)的
像素豁免压制。设计原则为"**色度证据只在模型不确定处拥有否决权**"——
背景色残留处模型原始输出低(实测中位 $\approx 0.28$),仍可压除;而背景色
的主体(浅蓝底上的白衬衫、粉底上的腮红)模型原始输出高(实测
$\approx 0.97$),被门保护。该门同时修复了纯 chroma 管线长期存在的
白衬衫/浅肤穿洞问题。
### 4.3 形态学清理
`clean_alpha` 移除未知带内面积 $< 10^{-5} HW$ 的孤立前景连通域;并将完全
落在 trimap-前景内部、面积 $\leq 2\times10^{-5} HW$ 且不触及图像边缘的
非实心区域填实(消除模型在确定前景内部的针孔)。
## 5. 前景颜色恢复与定向去溢色
### 5.1 前景色估计
以 pymatting 的多级前景估计(正则 $10^{-5}$)从 $I = \alpha F + (1-\alpha)B$
中解出 $F$,使半透明像素的颜色摆脱背景污染;不可用时回退启发式反混合。
### 5.2 定向去溢色(despill)
以背景 Lab 色度方向 $\hat d = (a_c, b_c)/\lVert\cdot\rVert$ 为轴,对前景带
(含向外扩张 2 px)中沿该轴超出边距 $m = 4.0$ 的色度分量按权回拉:
$$\mathrm{excess} = \max(\rho - m,\, 0), \qquad
(a, b) \leftarrow (a, b) - w \cdot \mathrm{excess} \cdot \hat d$$
权重 $w$ 由 alpha 权(不透明区全力,近透明区按下限 0.35 缓和)与色度权
(excess 归一)相乘。只动 $a^*b^*$、保留亮度 $L$,故发丝纹理与明暗不受损。
该式对任意平色背景成立:饱和绿幕退化为经典通道抑制,低饱和 pastel
背景(无主导 RGB 通道)由同一投影统一处理。
## 6. 设计原则小结
1. **分权与制衡**:拓扑归分割、颜色归色度、混合比例归 matting 模型;
任何一方不得越界(如色度不得推翻分割的拓扑判断)。
2. **证据分级**:证据越孤立,启用门槛越高(规则 2 的 $0.15$ 对规则 3 的
$0.60$)。
3. **弱者才可被否决**:后验修正(压制)只作用于上游不确定的区域——
trimap 限定作用域(未知带),matte 置信门限定作用强度。
4. **硬保证优先**:三值接口 + 硬钳制换来"确定区零泄漏"的结构保证;
连续证据经旁路(压制、置信门、alpha 种子)参与,不稀释保证。
5. **显式失败**:背景探测不达标即报错,拒绝在错误先验上静默产出。
## 7. 局限与评估中的扩展
**已知局限**:发丝间隙中颜色已漂移的背景残留(bgc $\approx 0.07$,偏离
$\tau_{fg}$ 之下)同时骗过三方——chroma 视其为前景、分割置信度高、
ViTMatte 判其为发丝高光($\alpha_{raw}$ 中位 $0.989$),四个模型族
(ViTMatte Com-1K / D646、BiRefNet_HR-matting、SEMat-SAM2)对其中约四成
像素一致判不透明。详见 docs/hair_gap_artifacts.md。
**交叉模型置信否决(已落地,默认开启)**:引入一个 trimap-free matting
模型(BiRefNet_HR-matting, 2048)作第二意见 $\alpha_{2nd}$,仅在色度定义
的可疑区($\rho \geq 3$ 且 $L \geq 45$ 且非确定背景,边界高斯羽化)内、
且主结果**假自信**的像素上行使否决,采用保守的 min 融合(视觉评审否决了
更激进的意见锐化变体):
$$\alpha' = \mathrm{lerp}\big(\alpha,\; \min(\alpha,\, \alpha_{2nd}),\; g\big),
\qquad g = \mathrm{zone} \times S(\alpha;\, 0.70,\, 0.95)$$
该设计延续原则 3:第二意见的发言范围由色度划定,可否决对象由置信门限定,
且只许压低、不许抬高——确定背景的零 alpha 结构上不可能被扰动;软发丝
(本就半透明,门控豁免)与暗发丝(亮度豁免)不受影响。可疑区内它有意
凌驾于 trimap-FG 硬钳制之上:待清除的残留多数恰是 trimap-FG。
## 附录:关键默认参数
| 模块 | 参数 | 默认值 |
|---|---|---|
| 背景探测 | 簇半径 / 边框占比门限 / 角容差 | 12 / 0.55 / 20 |
| 置信度场 | $\sigma_{\mathrm{lab}}$ 下限 / $\sigma_{\mathrm{rgb}}$ 下限 | 10 / 0.08 |
| Trimap | $\tau_{bg}$ / $\tau_{fg}$ / $\tau_{low}$ / $\tau_{core}$ / $\rho$ 门限 | 0.92 / 0.12 / 0.15 / 0.60 / 4.0 |
| 压制 | bgc 斜坡 / matte 门斜坡 / 强度 | 0.350.80 / 0.850.98 / 1.0 |
| 清理 | 最小连通域 / 填孔面积比 | $10^{-5}$ / $2\times10^{-5}$ |
| 去溢色 | 边距 $m$ / 强度 / alpha 权下限 | 4.0 / 0.92 / 0.35 |
| 交叉否决 | $\rho$ 门限 / $L$ 门限 / 羽化 $\sigma$ / 置信门 lohi | 3.0 / 45 / 2.0 / 0.700.95 |
| 模型 | 分割 / matting / 第二意见 / 设备 | BiRefNet@1024 / ViTMatte-base / BiRefNet_HR-matting@2048 / cpu |