True Virtual Try-On of Clothing with Latent Diffusion Models with Texture-Preserving Attention (TAPA)
Keywords:
Abstract
Virtual try-on (VTON) systems are designed to allow users to virtually try on clothes overlaid on their images, thus removing the need for physically trying on the clothes. Existing GAN-based VTON approaches face the perennial issue of inadequate garment texture preservation in cases of challenging body poses and occluded regions. Although latent diffusion models (LDMs) have advanced image synthesis fidelity recently, state-of-the-art diffusion-based VTON methods still suffer from significant loss of detail for patterned garments, for example, stripes, logos and printed textures. This work introduces a novel dual-branch LDM model with a TAPA module to enable garment-faithful VTON. We adopt a separate garment encoder to process the shop images of clothes and incorporate fine-grained texture features into the UNet in the denosing process through a cross-attention mechanism controlled by the pose (DensePose UV map) and segmentation maps. We also introduce an Appearance Consistency Loss (ACL) that operates in the Fourier frequency domain and penalties the discrepancy of texture in high-frequency range. Our experiments on the VITON-HD and Dress Code benchmark datasets show state-of-the-art result (SSIM 0.891, FID 7.15), improving upon the state-of-the-art baseline GarDiff by 1.1% and 15.2% with regard to SSIM and FID.
Downloads
Published
How to Cite
Issue
Section
License

This work is licensed under a Creative Commons Attribution-NoDerivatives 4.0 International License.