AViTS: Adaptive Spatiotemporal Token Selection for Efficient Dynamic-Resolution Generation
초록
Diffusion Transformers (DiTs) achieve high-quality generation but are costly due to iterative sampling. Dynamic-resolution sampling reduces early-stage cost by denoising at low resolution; however, uniformly upsampling all latent tokens at resolution transitions incurs redundant computation and may degrade fine-detail consistency. Existing partial upsampling strategies typically rely on local latent structure cues or single-step statistics, making it difficult to jointly capture token-text semantic relevance and token-wise representation dynamics across diffusion steps. We propose AViTS, an adaptive spatiotemporal token selection framework for dynamic-resolution DiTs. AViTS models spatial importance via latent-text attention and temporal importance via token-level feature variation across diffusion timesteps, and fuses them to enable spatiotemporal importance-aware selective upsampling: it prioritizes resolution refinement for critical tokens while deferring less important ones, thereby reducing redundant high-resolution computation and improving the quality-efficiency trade-off. AViTS achieves up to 6.34x on FLUX and nearly 9x FLOPs reduction on Qwen-Image-Edit and FLUX.1-Kontext-dev, orthogonal to distillation, quantization, and feature caching, and reaching 14.76x with distilled models. Code: https://github.com/QHR69/AViTS
저자 (12명)
- Haoran Qin — LinkedIn 검색
- Zhengan Yan — LinkedIn 검색
- Shikang Zheng — LinkedIn 검색
- Xiaobing Tu — LinkedIn 검색
- Jiacheng Liu — LinkedIn 검색
- Yuqi Lin — LinkedIn 검색
- Chang Zou — LinkedIn 검색
- JinShan Liu — LinkedIn 검색
- Peiliang Cai — LinkedIn 검색
- Xiantao Zhang — LinkedIn 검색
- Jinkui Ren — LinkedIn 검색
- Linfeng Zhang — LinkedIn 검색
저자 LinkedIn 변경 추적은 추후 자동화 예정입니다. 현재는 검색 링크를 제공합니다.