Files
lhk229 2b3b62ed3f Run the cross-check veto in complex-background mode with a hue-free gate
The veto's suspect zone was defined by background-hue projection, so
--no-chroma silently skipped it. Without a key colour there is no hue
test to lean on, so in complex mode the gate instead requires the second
opinion itself to be confidently near-empty (second_lo -> second_hi ramp,
default 0.15 -> 0.40): regions the HR-matting model decisively rejects
can be cleared, while thin strands it merely blurs to mid-alpha are
untouched -- protecting exactly the crisp-strand advantage the pipeline
has over a raw BiRefNet mask. Flat mode's gate is unchanged.

Validation (BG_IMAGE01-04, cuda fp32): flat default path bit-identical;
complex-mode veto touches 0.007-0.025% of pixels, visibly clearing
blurred residue near strands and milky specks in hair gaps with no
strand erosion. Costs the HR-matting forward in --no-chroma runs
(0.61 -> 1.6 s/image warm GPU); disable with --no-cross-check.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-14 13:09:57 +08:00

9.7 KiB

BgFilter

Offline character matting for AI-generated images on a flat-colour background. The background colour is auto-detected from the image border (green, pastel, any flat colour); pass --screen-color to set it explicitly.

RGB input
  -> chroma bg-confidence (keyed to the auto-detected or given colour)
  -> [optional] semantic segmentation mask
  -> trimap -> ViTMatte -> alpha cleanup
  -> cross-model veto (second matting opinion on bg-hued residue)
  -> pymatting foreground -> despill -> RGBA PNG -> QA previews

Pipelines

Two pipelines, selected by segmentation.enabled in the config:

  • Chroma-only (enabled: false) — Chroma + ViTMatte, no segmentation model. Lightest / fastest; leans entirely on the colour key for topology.
  • Single segmenter (enabled: true, default) — one segmentation model drives the trimap topology (holes, hair), ViTMatte then refines the soft edges. Backend is switchable: birefnet (default, general salient objects — text, logos, photos; needs trust_remote_code) or anime-seg (ONNX, tuned for anime characters). Switch at runtime with --seg-backend anime-seg (it also selects the matching weights).

Both share pymatting foreground estimation and an optional colour de-spill (off by default: de-spill has no positional/semantic guard, so a subject sharing the background's hue — a blue suit on a blue backdrop — gets desaturated and hue-shifted; enable with --despill when edge spill genuinely matters). There is no green-contamination rescue / recolour layer — with clean source images it is unnecessary, so it was removed.

Both also run a cross-model veto by default: a second, trimap-free matting model (ZhengPeng7/BiRefNet_HR-matting) may only lower alpha, only on background-hued bright pixels the primary result is confident about — this clears colour-drifted background residue trapped between hair strands that the chroma key, the segmenter and ViTMatte all read as foreground. Costs one extra model download (~0.9 GB) and one inference pass per image; disable with --no-cross-check (see the cross_check config section, and docs/hair_gap_artifacts.md for the analysis behind it).

When the cross-check is on, --cross-check-as-seg reuses that same HR-matting forward as the segmentation mask, skipping the primary seg model entirely (one less model to load, ~20 s faster per image on CPU). Off by default — use it only when the subjects are characters or similar solid figures. HR-matting is trained on portrait/animal matting (P3M-10k, AM-2k) and reads thin pale strokes on a low-contrast background as semi-transparent wisps: white calligraphy on the pastel backdrop (TestImage4) loses half its strokes under reuse, while the dedicated DIS5K-trained segmenter keeps them solid. The reuse is a same-family swap, so it applies to the birefnet backend only: with --seg-backend anime-seg the anime segmenter keeps its own forward and the cross-check veto still runs independently on top.

The segmentation trimap defaults to directional mode (chroma + seg + a hue-direction split: it keeps a background-coloured garment such as a white shirt while dropping a background-hued residual such as blue trapped between hair strands). Switch with --trimap-mode seg (topology only, no hue split) or directional-hard-bg (aggressive — hard-removes background-hued pixels; can eat cool/shadowed white cloth).

Non-flat backgrounds

Pass --no-chroma (or chroma.enabled: false) to matte images whose background is not one flat colour — gradients, textures, scenes. There is no colour key in this mode: the segmentation mask alone drives the trimap (mode forced to seg), ViTMatte still refines the unknown band at full resolution, and pymatting still estimates edge foreground colour. The colour-keyed stages are bypassed: background auto-detection, the directional hue split, chroma alpha suppression, and despill. The cross-model veto still runs (disable with --no-cross-check): without a key colour its suspect zone drops the hue test and instead only vetoes where the second opinion is itself confidently near-empty (cross_check.second_lo/hi), so thin strands the second model merely blurs are never eroded. Requires the segmentation pipeline; quality then rests entirely on the segmenter's mask, so expect flat-background results to stay stronger on hair-level detail.

Background colour

By default (screen_color: null) the background colour is auto-detected from the image border: the dominant flat colour of the border strip becomes the key colour. If the border is not one clean flat colour — a gradient, texture, or a subject filling the frame — detection fails with an error; pass --screen-color explicitly in that case.

To set it yourself, give a hex prior:

... --screen-color "#CFEFFF"

or screen_color: "#CFEFFF" in the config. Either way the chroma key scores pixels by perceptual (Lab/RGB) distance to the colour, and de-spill removes chroma along that colour's direction. (A supplied hex is refined against nearby border pixels; an auto-detected colour is used directly.)

Environment

Use the conda environment lightML.

conda activate lightML
pip install -r requirements.txt

If the shell is not activated, call the environment Python directly:

D:\MiniConda\envs\lightML\python.exe -m bgfilter.cli --help

Single Image

D:\MiniConda\envs\lightML\python.exe -m bgfilter.cli `
  --input Samples\TestImage.png `
  --output Outputs\TestImage_rgba.png `
  --debug-dir Outputs\TestImage_debug

No --config is needed — the built-in defaults are identical to configs/default.yaml. Pass --config configs\default.yaml only after you edit that file to tune the detailed parameters. Runtime choices stay on the command line: --device (default CPU; use --device cuda for GPU), --precision (default fp32; bf16 speeds up all three models and halves the matting model's activation memory with visually identical alpha — needs bf16-capable hardware, falls back to fp32 elsewhere; large inputs additionally get ViTMatte's global attention computed in query chunks by default — exact, bitwise-identical, caps the memory spike at ~4 GB instead of ~19 GB at 2048x2048, see model.attn_query_chunk), --seg-backend (default birefnet; anime-seg for anime characters), --screen-color (default: auto-detect the flat background), and --trimap-mode.

Batch

D:\MiniConda\envs\lightML\python.exe -m bgfilter.cli `
  --input-dir Samples `
  --output-dir Outputs `
  --debug-dir Outputs\debug

Performance (CPU reference numbers)

Measured on a Ryzen 9700X (Zen 5, native bf16), 32 GB RAM, with cross-check on and --cross-check-as-seg (reuse); chunked global attention on. Without reuse (the default) add ~9 s fp32 / ~6 s bf16 for the dedicated seg forward:

input precision warm / image peak memory
1024x1536 fp32 ~45 s
1024x1536 bf16 ~33 s
2048x2048 bf16 ~53 s 11.4 GB (batch), 8.1 GB (single image)

The memory ceiling is the cross-check BiRefNet_HR forward (fixed input_size 2048 regardless of the image size) — genuine live activations of a full-resolution dense prediction net. Before chunked attention and the mimalloc fix, the same 2048x2048 bf16 run peaked at 25.2 GB; a fully un-optimized fp32 run would need an estimated 55-60 GB (ViTMatte's un-chunked N^2 attention alone ~39 GB). CPUs without native bf16 (e.g. Zen 2) auto-fall back to fp32 — on such machines prefer --no-cross-check if memory or latency is tight.

Chroma-alpha debug mode

--matting-method chroma skips ViTMatte and uses chroma confidence directly as the alpha seed. Useful for fast inspection of chroma confidence, trimap, and despill. (Distinct from the chroma-only pipeline above, which still runs ViTMatte.)

D:\MiniConda\envs\lightML\python.exe -m bgfilter.cli `
  --input-dir Samples `
  --output-dir Outputs\chroma `
  --debug-dir Outputs\chroma_debug `
  --matting-method chroma `
  --device cpu

Outputs

For each processed image, the CLI writes an RGBA PNG and optional debug files:

bg_confidence.png
trimap.png
seg_mask.png          # segmentation pipeline only
alpha.png
foreground_rgb.png    # despilled foreground colour
color_mask.png        # per-pixel despill weight
preview_black.png
preview_white.png
preview_gray.png
preview_red.png
preview_blue.png
qa_grid.png
metadata.json

Quality Check

The quality checker measures alpha validity and edge spill on semi-transparent edge pixels.

D:\MiniConda\envs\lightML\python.exe -m bgfilter.quality_cli `
  Outputs\TestImage_rgba.png `
  --max-edge-green-excess-p95 0.30

Run the bundled sample smoke check:

D:\MiniConda\envs\lightML\python.exe scripts\smoke_samples.py `
  --samples-dir Samples `
  --output-dir Outputs\smoke_samples `
  --config configs\default.yaml `
  --device cpu `
  --fallback-to-chroma-alpha `
  --max-edge-green-excess-p95 0.30

Notes

  • Samples/ and Outputs/ are ignored by Git.
  • ViTMatte and segmentation weights load from Hugging Face on first use. Behind a firewall set HF_ENDPOINT=https://hf-mirror.com (and bypass a flaky local proxy). anime-seg (skytnt/anime-seg) is a plain ONNX download; birefnet (ZhengPeng7/BiRefNet) ships custom modelling code so it needs trust_remote_code=True plus timm / einops / kornia.
  • Foreground colour estimation uses pymatting's estimate_foreground_ml to propagate clean foreground colour into semi-transparent edges before de-spill. Set foreground.method: unmix to fall back to the legacy heuristic.
  • docs/green_screen_matting_workflow.md is the original phase-1 green-screen spec; this README reflects the current, generalised architecture.