Score-Scaled Residual-Site Steering in Multimodal Diffusion Transformers

Ibragim Idrisov, Dmitry Namiot
15m
We study inference-time vector steering for multimodal diffusion transformers. Unlike U-Net diffusion models, recent text-to-image generators expose no single cross-attention interface for editing text influence. We formulate steering as residual-state transport inside transformer blocks and introduce a score-scaled mean transport estimator. The method estimates concept displacements from contrastive prompt pairs and rescales them with two score masks: a reliability mask measuring prompt-pair consistency and a robust effect mask measuring the displacement in median-absolute-deviation units. The resulting intervention preserves the empirical source-to-target magnitude while suppressing unstable or scale-dominated coordinates. We evaluate object, attribute, and style steering with VQA-based target success, CLIP prompt preference, and DreamSim preservation.