Virtual try-on is image generation with a strict brief: change one thing, keep everything else.
A try-on model must change exactly one thing, the clothing, and keep everything else: the person's face and identity, their pose and body, the light and the background. Unconstrained generators are very good at making plausible images. The hard part is making the right one, with the exact garment from the product page.
The field moved from warping the product image with an explicit geometric transform and blending it in, to diffusion models that learn where each part of the garment should go through attention. Diffusion brought large gains in realism and in robustness to pose. It also brought new problems: keeping fine detail through a compressed latent space, keeping identity exact, and sampling fast enough for someone who is shopping right now.
Questions we are exploring
- How do we carry small, high-frequency detail such as text and logos through a latent diffusion model?
- How do we keep a person's identity exact while changing what they wear?
- How few sampling steps can we use before quality drops where people actually look?
Further reading
- Zhu et al. TryOnDiffusion: A Tale of Two UNets. CVPR, 2023.
- Kim et al. StableVITON: Learning Semantic Correspondence with Latent Diffusion Model for Virtual Try-On. CVPR, 2024.
- Choi et al. Improving Diffusion Models for Authentic Virtual Try-on in the Wild. ECCV, 2024.
- Xu et al. OOTDiffusion: Outfitting Fusion based Latent Diffusion for Controllable Virtual Try-on. AAAI, 2025.
- Chong et al. CatVTON: Concatenation Is All You Need for Virtual Try-On with Diffusion Models. ICLR, 2025.
- Luo et al. Latent Consistency Models: Synthesizing High-Resolution Images with Few-Step Inference. arXiv, 2023.
- Sauer et al. Adversarial Diffusion Distillation. ECCV, 2024.


