2cc0ce8684
ViTMatte's VitDet backbone runs 4 global attention blocks that materialize the full [heads x N x N] map: ~19 GB transient at 2048x2048 (16384 tokens), the pipeline's memory peak. transformers has no SDPA path for this architecture (the decomposed rel-pos bias is added to raw scores), so compute the same attention in query-row chunks instead: the bias factorizes over query rows, making the chunked form mathematically exact -- output verified BITWISE-identical (unit: fp32/bf16 x 3 sizes; full pipeline: TestImage3 fp32 and 2048x2048 bf16, all byte-equal). model.attn_query_chunk (default 2048, 0 = stock one-shot) engages only on blocks seeing more tokens than the chunk size, so window blocks keep the original path. Measured @2048x2048 bf16 (9700X): ViTMatte spike 19.8 -> 4.0 GB for ~15% more ViTMatte time; whole pipeline peak 21.4 -> 11.4 GB, warm 51 -> 53 s. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>