# BgFilter Offline character matting for AI-generated images on a flat-colour background. The background colour is **auto-detected** from the image border (green, pastel, any flat colour); pass `--screen-color` to set it explicitly. ```text RGB input -> chroma bg-confidence (keyed to the auto-detected or given colour) -> [optional] semantic segmentation mask -> trimap -> ViTMatte -> alpha cleanup -> cross-model veto (second matting opinion on bg-hued residue) -> pymatting foreground -> despill -> RGBA PNG -> QA previews ``` ## Pipelines Two pipelines, selected by `segmentation.enabled` in the config: - **Chroma-only** (`enabled: false`) — Chroma + ViTMatte, no segmentation model. Lightest / fastest; leans entirely on the colour key for topology. - **Single segmenter** (`enabled: true`, default) — one segmentation model drives the trimap topology (holes, hair), ViTMatte then refines the soft edges. Backend is switchable: `birefnet` (default, general salient objects — text, logos, photos; needs `trust_remote_code`) or `anime-seg` (ONNX, tuned for anime characters). Switch at runtime with `--seg-backend anime-seg` (it also selects the matching weights). Both share pymatting foreground estimation and an optional colour de-spill (**off by default**: de-spill has no positional/semantic guard, so a subject sharing the background's hue — a blue suit on a blue backdrop — gets desaturated and hue-shifted; enable with `--despill` when edge spill genuinely matters). There is **no** green-contamination rescue / recolour layer — with clean source images it is unnecessary, so it was removed. Both also run a **cross-model veto** by default: a second, trimap-free matting model (`ZhengPeng7/BiRefNet_HR-matting`) may only *lower* alpha, only on background-hued bright pixels the primary result is confident about — this clears colour-drifted background residue trapped between hair strands that the chroma key, the segmenter and ViTMatte all read as foreground. Costs one extra model download (~0.9 GB) and one inference pass per image; disable with `--no-cross-check` (see the `cross_check` config section, and `docs/hair_gap_artifacts.md` for the analysis behind it). When the cross-check is on, `--cross-check-as-seg` reuses that same HR-matting forward as the segmentation mask, skipping the primary seg model entirely (one less model to load, ~20 s faster per image on CPU). **Off by default** — use it only when the subjects are characters or similar solid figures. HR-matting is trained on portrait/animal matting (P3M-10k, AM-2k) and reads thin pale strokes on a low-contrast background as semi-transparent wisps: white calligraphy on the pastel backdrop (TestImage4) loses half its strokes under reuse, while the dedicated DIS5K-trained segmenter keeps them solid. The reuse is a same-family swap, so it applies to the `birefnet` backend only: with `--seg-backend anime-seg` the anime segmenter keeps its own forward and the cross-check veto still runs independently on top. The segmentation trimap defaults to `directional` mode (chroma + seg + a hue-direction split: it keeps a background-coloured garment such as a white shirt while dropping a background-hued residual such as blue trapped between hair strands). Switch with `--trimap-mode seg` (topology only, no hue split) or `directional-hard-bg` (aggressive — hard-removes background-hued pixels; can eat cool/shadowed white cloth). ## Non-flat backgrounds Pass `--no-chroma` (or `chroma.enabled: false`) to matte images whose background is **not** one flat colour — gradients, textures, scenes. There is no colour key in this mode: the segmentation mask alone drives the trimap (mode forced to `seg`), ViTMatte still refines the unknown band at full resolution, and pymatting still estimates edge foreground colour. The colour-keyed stages are bypassed: background auto-detection, the directional hue split, chroma alpha suppression, and despill. The cross-model veto still runs (disable with `--no-cross-check`): without a key colour its suspect zone drops the hue test and instead only vetoes where the second opinion is itself confidently near-empty (`cross_check.second_lo/hi`), so thin strands the second model merely blurs are never eroded. Requires the segmentation pipeline; quality then rests entirely on the segmenter's mask, so expect flat-background results to stay stronger on hair-level detail. ## Background colour By default (`screen_color: null`) the background colour is **auto-detected** from the image border: the dominant flat colour of the border strip becomes the key colour. If the border is not one clean flat colour — a gradient, texture, or a subject filling the frame — detection **fails with an error**; pass `--screen-color` explicitly in that case. To set it yourself, give a hex prior: ```powershell ... --screen-color "#CFEFFF" ``` or `screen_color: "#CFEFFF"` in the config. Either way the chroma key scores pixels by perceptual (Lab/RGB) distance to the colour, and de-spill removes chroma along that colour's direction. (A supplied hex is refined against nearby border pixels; an auto-detected colour is used directly.) ## Environment Use the conda environment `lightML`. ```powershell conda activate lightML pip install -r requirements.txt ``` If the shell is not activated, call the environment Python directly: ```powershell D:\MiniConda\envs\lightML\python.exe -m bgfilter.cli --help ``` ## Single Image ```powershell D:\MiniConda\envs\lightML\python.exe -m bgfilter.cli ` --input Samples\TestImage.png ` --output Outputs\TestImage_rgba.png ` --debug-dir Outputs\TestImage_debug ``` No `--config` is needed — the built-in defaults are identical to `configs/default.yaml`. Pass `--config configs\default.yaml` only after you edit that file to tune the detailed parameters. Runtime choices stay on the command line: `--device` (default **CPU**; use `--device cuda` for GPU), `--precision` (default `fp32`; `bf16` speeds up all three models and halves the matting model's activation memory with visually identical alpha — needs bf16-capable hardware, falls back to fp32 elsewhere; large inputs additionally get ViTMatte's global attention computed in query chunks by default — exact, bitwise-identical, caps the memory spike at ~4 GB instead of ~19 GB at 2048x2048, see `model.attn_query_chunk`), `--seg-backend` (default `birefnet`; `anime-seg` for anime characters), `--screen-color` (default: auto-detect the flat background), and `--trimap-mode`. ## Batch ```powershell D:\MiniConda\envs\lightML\python.exe -m bgfilter.cli ` --input-dir Samples ` --output-dir Outputs ` --debug-dir Outputs\debug ``` ## Performance (CPU reference numbers) Measured on a Ryzen 9700X (Zen 5, native bf16), 32 GB RAM, with cross-check on and `--cross-check-as-seg` (reuse); chunked global attention on. Without reuse (the default) add ~9 s fp32 / ~6 s bf16 for the dedicated seg forward: | input | precision | warm / image | peak memory | |---|---|---|---| | 1024x1536 | fp32 | ~45 s | — | | 1024x1536 | bf16 | ~33 s | — | | 2048x2048 | bf16 | ~53 s | 11.4 GB (batch), 8.1 GB (single image) | The memory ceiling is the cross-check BiRefNet_HR forward (fixed `input_size` 2048 regardless of the image size) — genuine live activations of a full-resolution dense prediction net. Before chunked attention and the mimalloc fix, the same 2048x2048 bf16 run peaked at 25.2 GB; a fully un-optimized fp32 run would need an estimated 55-60 GB (ViTMatte's un-chunked N^2 attention alone ~39 GB). CPUs without native bf16 (e.g. Zen 2) auto-fall back to fp32 — on such machines prefer `--no-cross-check` if memory or latency is tight. ## Chroma-alpha debug mode `--matting-method chroma` skips ViTMatte and uses chroma confidence directly as the alpha seed. Useful for fast inspection of chroma confidence, trimap, and despill. (Distinct from the chroma-only *pipeline* above, which still runs ViTMatte.) ```powershell D:\MiniConda\envs\lightML\python.exe -m bgfilter.cli ` --input-dir Samples ` --output-dir Outputs\chroma ` --debug-dir Outputs\chroma_debug ` --matting-method chroma ` --device cpu ``` ## Outputs For each processed image, the CLI writes an RGBA PNG and optional debug files: ```text bg_confidence.png trimap.png seg_mask.png # segmentation pipeline only alpha.png foreground_rgb.png # despilled foreground colour color_mask.png # per-pixel despill weight preview_black.png preview_white.png preview_gray.png preview_red.png preview_blue.png qa_grid.png metadata.json ``` ## Quality Check The quality checker measures alpha validity and edge spill on semi-transparent edge pixels. ```powershell D:\MiniConda\envs\lightML\python.exe -m bgfilter.quality_cli ` Outputs\TestImage_rgba.png ` --max-edge-green-excess-p95 0.30 ``` Run the bundled sample smoke check: ```powershell D:\MiniConda\envs\lightML\python.exe scripts\smoke_samples.py ` --samples-dir Samples ` --output-dir Outputs\smoke_samples ` --config configs\default.yaml ` --device cpu ` --fallback-to-chroma-alpha ` --max-edge-green-excess-p95 0.30 ``` ## Notes - `Samples/` and `Outputs/` are ignored by Git. - ViTMatte and segmentation weights load from Hugging Face on first use. Behind a firewall set `HF_ENDPOINT=https://hf-mirror.com` (and bypass a flaky local proxy). `anime-seg` (`skytnt/anime-seg`) is a plain ONNX download; `birefnet` (`ZhengPeng7/BiRefNet`) ships custom modelling code so it needs `trust_remote_code=True` plus `timm` / `einops` / `kornia`. - Foreground colour estimation uses pymatting's `estimate_foreground_ml` to propagate clean foreground colour into semi-transparent edges before de-spill. Set `foreground.method: unmix` to fall back to the legacy heuristic. - `docs/green_screen_matting_workflow.md` is the original phase-1 green-screen spec; this README reflects the current, generalised architecture.