MRI
MRI India Journals Vol. 13 No. 2S (2026): Special Issue: ICSAIEM

True Virtual Try-On of Clothing with Latent Diffusion Models with Texture-Preserving Attention (TAPA)

Authors

  • Aashutosh Chandgude Department of AI & DS, Dr. D. Y. Patil College of Engineering & Innovation, Pune, India.
  • Maithili Bhosale Department of AI & DS, Dr. D. Y. Patil College of Engineering & Innovation, Pune, India.
  • Riddhesh Jadav Department of AI & DS, Dr. D. Y. Patil College of Engineering & Innovation, Pune, India.
  • Sairaj Pawar Department of AI & DS, Dr. D. Y. Patil College of Engineering & Innovation, Pune, India.

Keywords:

Virtual Try-On Latent Diffusion Models Garment Texture Preservation Cross-Attention VITON-HD Computer Vision Generative AI Deep Learning

Abstract

Virtual try-on (VTON) systems are designed to allow users to virtually try on clothes overlaid on their images, thus removing the need for physically trying on the clothes. Existing GAN-based VTON approaches face the perennial issue of inadequate garment texture preservation in cases of challenging body poses and occluded regions. Although latent diffusion models (LDMs) have advanced image synthesis fidelity recently, state-of-the-art diffusion-based VTON methods still suffer from significant loss of detail for patterned garments, for example, stripes, logos and printed textures. This work introduces a novel dual-branch LDM model with a TAPA module to enable garment-faithful VTON. We adopt a separate garment encoder to process the shop images of clothes and incorporate fine-grained texture features into the UNet in the denosing process through a cross-attention mechanism controlled by the pose (DensePose UV map) and segmentation maps. We also introduce an Appearance Consistency Loss (ACL) that operates in the Fourier frequency domain and penalties the discrepancy of texture in high-frequency range. Our experiments on the VITON-HD and Dress Code benchmark datasets show state-of-the-art result (SSIM 0.891, FID 7.15), improving upon the state-of-the-art baseline GarDiff by 1.1% and 15.2% with regard to SSIM and FID.

Downloads

Published

2026-07-05

How to Cite

Chandgude, A., Bhosale, M., Jadav, R., & Pawar, S. (2026). True Virtual Try-On of Clothing with Latent Diffusion Models with Texture-Preserving Attention (TAPA). Multidisciplinary Journal of Research in Engineering and Technology, 13(2S), 319–324. Retrieved from https://journals.mriindia.com/index.php/mjret/article/view/4035

Similar Articles

<< < 5 6 7 8 9 10 11 12 13 14 > >> 

You may also start an advanced similarity search for this article.