Commit Graph

12 Commits

Author SHA1 Message Date
lhk229 c56e58affe Add complex-background matting mode to the service via background_mode
Backports master's non-flat matting (chroma.enabled: false + the hue-free
cross-check gate) into server-edition, and exposes it over HTTP without
surfacing the internal "chroma" wording: /remove-background gains a
background_mode form field (flat, default | complex). complex maps to
chroma disabled -- no colour key, segmentation alone drives the trimap and
every colour-keyed stage (auto-detect, hue split, chroma suppression,
despill) is bypassed. The cross-check veto still works in complex mode via
its second-opinion-confidence gate (cross_check.second_lo/hi) but defaults
OFF there (it costs the HR-matting forward); an explicit cross_check=on
re-enables it.

No new model weights: complex mode reuses the already-provisioned BiRefNet
seg + ViTMatte (+ optional HR-matting cross-check). Flat mode is unchanged
(bit-identical), and server-edition's own extras (cross_check.lock,
foreground.use_gpu CuPy path) are preserved -- the port is surgical, not a
copy of master's files.

- settings: ChromaSettings.enabled, CrossCheckSettings.second_lo/hi
- config: override_settings chroma passthrough
- despill/foreground: model=None safe guards (foreground keeps GPU path)
- alpha_post: cross_check_alpha hue-free gate when proj is None
- pipeline: _process_rgb complex branch (seg-only trimap, skip colour stages)
- service/app: process(chroma=), background_mode field, complex-defaults-off
  cross-check, X-BGFilter-Background-Mode header
- cli: --chroma/--no-chroma
- configs/docs: gpu.yaml + default.yaml + README/README_ZH

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-14 14:04:58 +08:00
lhk229 9b9d71cf82 Default reuse-as-seg OFF; default expandable_segments on Linux/WSL
Two default flips for server-edition:

1. cross_check.reuse_as_seg now defaults to false (settings + default.yaml + CLI help): the dedicated segmenter keeps its own forward and the cross-check veto stays an independent second signal. The reuse remains available via --cross-check-as-seg / config; measured cost of off vs on: ~20 s/image on CPU, +0.1-0.25 s and +0.9-1.3 GB VRAM on GPU (13.3 vs 12.0 GB reserved under expandable segments -- still fits a 16 GB card).

2. bgfilter/__init__.py defaults PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True on non-Windows platforms, before torch loads (setdefault: an explicit env value wins; native Windows is excluded because torch warns and ignores it there). Measured: reserved 15.3 -> 12.0 GB and ~10% faster on the RTX 5070 Ti; no effect on CPU-only runs.

README/DEPLOY(_ZH) synced: reuse documented as opt-in, GPU profile numbers updated for both reuse states, perf table labeled with the config it was measured under.

Verified: defaults resolve off/set as intended on Windows and WSL; explicit PYTORCH_CUDA_ALLOC_CONF override wins.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-07 20:40:19 +08:00
lhk229 3992f65ebf Return freed activations to the OS on Linux; recommend jemalloc (A+B)
MIMALLOC_PURGE_DELAY=0 only bites on Windows (bundled mimalloc). On a Linux
server PyTorch uses glibc ptmalloc, which keeps a BiRefNet@2048 forward's freed
activations in the arena, so RSS ratchets up across requests and the env var is
a no-op there.

A. bgfilter/memtune.py: release_freed_memory() calls glibc malloc_trim(0) at
   runtime (no-op on Windows/musl/other allocators). Wired into service.process
   (after each request) and the CLI batch loop (after each image), the two
   long-lived paths where RSS accumulates. Single-image CLI exits, so it is left
   alone.
B. DEPLOY(_ZH): document preloading jemalloc via LD_PRELOAD + MALLOC_CONF (also
   improves CPU throughput) as the production alternative, with MALLOC_ARENA_MAX/
   MALLOC_TRIM_THRESHOLD_ as an allocator-free fallback. Under jemalloc/tcmalloc
   malloc_trim simply no-ops.

Also fixes the stale off-by-default cross-check heading in DEPLOY.md.

Verified: modules import; release_freed_memory() returns False (no-op) on Windows.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-06 16:40:58 +08:00
lhk229 9a9084b3d6 Cut pipeline cost: seg-reuse, unified bf16 knob, eager mimalloc purge
Three optimizations from profiling the cross-check-dominated pipeline
(9700X CPU, all pilot-validated on TestImage3/FixImage1):

- Reuse the cross-check HR-matting@2048 forward as the segmentation mask
  (cross_check.reuse_as_seg, default ON; --no-cross-check-as-seg to opt
  out). Skips the BiRefNet@1024 load+forward entirely: ~66s -> ~45s,
  one less 0.9GB model. Trimap 99.8% identical, no structural change.

- --precision bf16 now fans out to all three models: ViTMatte keeps its
  weight cast; both BiRefNets run their forward under autocast with a
  dispatcher-level AutocastCPU fp32 shim for torchvision::deform_conv2d
  (no bf16 CPU kernel, no autocast wrapper upstream). Shared hardware
  gate in bgfilter/precision.py falls back to fp32 off native-bf16
  hardware. TestImage3: 51.9s -> 33.1s; alpha diff max 0.15, none >0.25.

- MIMALLOC_PURGE_DELAY=0 (bgfilter/__init__.py, before torch loads):
  Windows torch's bundled mimalloc lazily retains ~10GB of freed
  BiRefNet activations, stacking under ViTMatte's attention peak.
  2048x2048 bf16: peak 25.2 -> 21.4GB and slightly faster (73 -> 63s).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-06 15:15:12 +08:00
lhk229 dfc85566c7 Add bf16 precision option for ViTMatte (--precision, default fp32)
bf16 halves the matting model's activation memory -- its full-resolution
attention is the pipeline's memory peak -- with visually identical alpha
(measured: 0 px alpha deviation > 0.25 on samples; cross-check veto
behaviour unchanged, region IoU 0.984).

Guarded by a hardware check so it never lands on a slow emulation path:
CUDA requires torch.cuda.is_bf16_supported(); CPU requires the same
oneDNN native-bf16 gate PyTorch uses for matmul routing (AVX512-BF16/
AMX). Without support it warns and falls back to fp32 -- on a Zen2 EPYC
the fallback kernels measured 17-370x slower than fp32, so silent bf16
there would be a performance landmine.

Segmentation stays fp32: torchvision deform_conv2d (used by BiRefNet)
has no bf16 CPU kernel, and the segmenter is not the memory peak.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-06 13:32:51 +08:00
lhk229 adbbb9d620 Add cross-model veto of hair-gap background residue (default on)
Colour-drifted background trapped between hair strands defeats every
single-signal defence: the chroma key reads it as foreground (bgc ~0.07),
the segmenter backs it, ViTMatte rates it opaque, and post-hoc removal is
a proven dead end (it shreds the hair volume the same pixels belong to).
A second, trimap-free matting model (BiRefNet_HR-matting) is the only
tested model that separates this residue from the subject, so its opinion
is fused in as a veto: min-fusion that may only LOWER alpha, restricted to
the background-hued bright suspect zone (proj >= 3, L >= 45, feathered)
and gated by primary-alpha confidence (0.70 -> 0.95 ramp) so soft wisps
and dark hair are exempt by construction.

- settings/config/CLI: cross_check block, --cross-check/--no-cross-check
- alpha_post.cross_check_alpha after clean_alpha; second opinion reuses
  BiRefNetSegmenter; saved to debug as cross_check_alpha.png
- chroma.bg_hue_projection extracted and shared with the trimap
- docs: methodology.md (new), hair_gap_artifacts.md (investigation log)

Verified: cross-check ON reproduces the visually-reviewed B1gate
prototype byte-for-byte on TestImage3; --no-cross-check reproduces the
previous baseline byte-for-byte; pink-bg FixImage1 face untouched.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-04 19:52:24 +08:00
lhk229 a6d84845c6 Default to BiRefNet + CPU; add --seg-backend to pick the segmenter
- Segmentation backend now defaults to birefnet (ZhengPeng7/BiRefNet); anime-seg is
  selected at runtime with --seg-backend anime-seg, which also swaps in the matching
  weights (unknown backends raise). Config files stay for detailed tuning, not
  backend selection.
- Default device is now cpu everywhere (dataclass defaults, default.yaml,
  birefnet.yaml); pass --device cuda for GPU. Avoids a hard CUDA-required failure on
  machines without a CUDA-enabled PyTorch build.
- Docs: README + workflow status note updated (backend via --seg-backend, CPU
  default, examples no longer force --device cuda).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-02 16:32:13 +08:00
lhk229 62f175948e Auto-detect the background colour (default); drop the green heuristic
Replace the hardcoded green-screen auto-detection with a general flat-colour
detector. When screen_color is null the background colour is now detected from the
image border (dominant colour of the border strip) and used directly as the chroma
model, with a failure gate that raises when the border is not one clean flat colour
(gradient / texture / subject filling the frame).

- chroma: add detect_background_color + _background_border_cluster; the null case of
  estimate_background_model samples the detected cluster directly (no seed search);
  compute_bg_confidence is now always perceptual (Lab/RGB distance). Removed the green
  heuristic (initial_green_candidates), the green multiplicative gating, the now-unused
  _smoothstep, and five green-only ChromaSettings fields.
- pipeline: reuse the detected colour for de-spill in the auto case.
- settings / default.yaml: add detect_* tuning fields; refresh screen_color docs.
- README: document auto-detection + the failure gate, trimap modes, birefnet config.

A supplied --screen-color hex still uses the seed-search refinement. Validated:
detects green / pastel / text samples correctly, raises on gradient / two-colour, CLI
exits 1 on failure, green and pastel mattes unchanged in quality.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-01 20:34:31 +08:00
lhk229 5f48f7a2cd Add directional trimap mode (default); keep seg/hard-bg via trimap.mode
Replace the two experimental bools (directional, bg_hued_to_bg) with a single
TrimapSettings.mode selector for the segmentation pipeline:
  - "directional" (new default): chroma magnitude + seg + a Lab hue-direction
    split. In the chroma-unknown zone a confidently-segmented pixel stays
    foreground unless it is displaced toward the background hue, so a neutral
    background-coloured garment (e.g. a white shirt) is kept while a background-
    hued residual (blue between hair strands) is left unknown for ViTMatte /
    chroma-suppress to clear.
  - "seg": original fuse_trimap, unchanged.
  - "directional-hard-bg": aggressive variant that hard-removes background-hued
    pixels (can eat cool/shadowed white cloth).

Selectable via configs/default.yaml (trimap.mode) or CLI --trimap-mode; unknown
modes raise. Default CLI output verified byte-identical to the reviewed UNK
result on the pastel sample.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-01 19:27:09 +08:00
lhk229 7555321cfc Simplify to single-segmenter pipelines; add screen_color prior; drop rescue + recolor
Strategic pivot: the dropout-fill rescue, the anime-seg fill backend, the
BiRefNet+anime-seg intersection, and the recolor pass all existed to repair
source images whose hair was already green-contaminated or broken at generation
time. Fixing the source instead (clean pastel-background generation) makes the
matte high-contrast, so that whole compensation layer is unnecessary. Remove it.

- Two pipelines via segmentation.enabled:
    false -> chroma-only (Chroma + ViTMatte)
    true  -> single segmenter (default anime-seg, switchable birefnet) -> trimap
  then ViTMatte refines, pymatting estimates foreground, despill cleans spill.

- Generalise the green-hardcoded colour logic to a screen_color prior (hex, e.g.
  "#CFEFFF"; default null = green auto-detect, byte-identical). chroma keys off
  Lab/RGB distance to the colour; despill removes chroma along its Lab direction.
  Added --screen-color CLI flag.

- Removed dead code: fill_seg_dropouts + seg_fill*/intersection params,
  SegmentationSettings.fill_backend/fill_model_name/sharpen, unsharp_mask,
  recolor.py + RecolorSettings.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-01 15:33:58 +08:00
Codex 10705d49df Support YAML pipeline configuration 2026-06-30 15:03:15 +08:00
Codex 8fc53925e0 Implement green screen matting pipeline 2026-06-30 14:14:15 +08:00