Skip to content
Clothsy AI Talk to us

Field notes9 min readFabricVTON Research

A short history of virtual try-on

From warping a product photo onto a body to diffusion models that learn where every thread should go: how the field got here, and what is still unsolved.

Virtual try-on asks a question that sounds simple: given a photo of a person and a photo of a garment, what would that person look like wearing it? In its modern, image-based form the research field is less than a decade old. In that time it has already been through one complete change of method, from explicitly warping product images onto bodies to diffusion models that learn the correspondence for themselves. This is a short tour of how it got here, and of what is still open.

The task

A try-on system receives two images and must produce a third. From the person photo it has to keep almost everything: face and identity, pose, body shape, skin, hair and the background. From the product photo it has to take the garment: its shape, colour, fabric, print and details. The output has to look like a real photograph, and it has to be faithful to both inputs at once.

Person photo (input)
Person photo
Product photo of a garment (input)
Product photo
Generated try-on result (output)
Generated try-on

Keep from the personFace and identity, pose, body shape, skin, hair, background.

Take from the productGarment shape, colour, fabric texture, print, logos, details like buttons and seams.

Figure 1The task. A person photo and a product photo go in; a new image of that person wearing that product comes out. Images: the Clothsy AI demo set.

One idea runs through almost every system since the first: remove the old clothing from the person before adding the new. VITON [1] called this a clothing-agnostic person representation. It keeps pose keypoints, a coarse body-shape mask and the face and hair, and throws away what the person was wearing, so that the model cannot simply copy it back.

Warp, then blend (2018–2022)

The first generation of systems split the problem in two: geometrically deform the product image to fit the body, then blend it into the person. VITON produced a coarse result first, then warped the garment with a thin-plate spline (TPS), a smooth deformation controlled by a grid of points, and used a refinement network to blend the warped garment into the coarse image [1]. Its warp came from classical shape matching. CP-VTON made the warp learnable, with a Geometric Matching Module that predicts the TPS transform, and a try-on module that learns a composition mask deciding, pixel by pixel, how much of the warped cloth to use [2].

Resolution and robustness came next. VITON-HD generated images at 1024×768 and introduced a normalization designed for the regions where the warped garment and the body do not line up [3]. HR-VITON performed warping and segmentation in a single module, to handle misalignment and the occlusions that make a garment look squeezed where the body covers it [4].

Explicit warpTPSMove a control grid, warp the productimage, then blend it into the person.Implicit correspondenceEach output region looks up thegarment features it needs.
Figure 2Two ways to move a garment onto a body. Older systems warp the product image with an explicit transform, then blend it in. Diffusion systems learn the correspondence inside the network, through attention.

Explicit warping has a real strength. When the pose is close to the product photo, it carries texture across almost untouched, because it is literally moving the original pixels. Its weakness is everything a single smooth warp cannot express: large pose changes, arms crossing the body, loose garments, layers. And the blending networks of the time, mostly GANs, left visible seams and artifacts when the warp was wrong.

Generate with attention (2023 onwards)

Diffusion models changed the approach. TryOnDiffusion used two UNets working in parallel, one for the person and one for the garment, and let cross-attention warp the garment implicitly. There is no explicit geometric transform; the network learns which part of the garment belongs where [5]. It handled large changes in pose and body shape, and a cascade of diffusion models brought the output to 1024×1024.

A burst of work followed, mostly built on large pre-trained latent diffusion models:

  • LaDI-VTON added textual inversion, mapping garment features into the text-token space the diffusion model already understands [6].
  • StableVITON learned the correspondence between clothing and body with zero-initialized cross-attention blocks inside a pre-trained model’s latent space [7].
  • IDM-VTON split the garment signal in two: high-level semantics through an image-prompt adapter, and low-level detail through a parallel garment UNet [8].
  • OOTDiffusion learned garment features in an outfitting UNet and fused them into the denoising network through self-attention [9].
  • CatVTON went the other way on complexity: it simply concatenates the garment and person images side by side, with no extra image encoder, and trains only a small fraction of the model’s parameters [10].
  • M&M VTO moved beyond single garments to full outfits, with the layout controlled by text [11].

Warp, then blendGAN era

  1. 2018VITONClothing-agnostic person, TPS warp, refinement
  2. 2018CP-VTONLearned geometric matching, composition mask
  3. 2021VITON-HD1024×768, misalignment-aware normalization
  4. 2022HR-VITONWarping and segmentation in one module

Generate with attentionDiffusion era

  1. 2023TryOnDiffusionParallel-UNet, implicit warp by cross-attention
  2. 2023LaDI-VTONLatent diffusion with textual inversion
  3. 2024StableVITONZero cross-attention correspondence
  4. 2024IDM-VTONGarment UNet plus image-prompt adapter
  5. 2024M&M VTOSeveral garments, layout by text
  6. 2025OOTDiffusionOutfitting UNet and fusion
  7. 2025CatVTONGarment and person concatenated
Figure 3A selection of influential papers, by year of the peer-reviewed venue. The field moved from explicit warping with GANs to diffusion models that learn the correspondence themselves.

Diffusion brought large gains in realism and in robustness to pose. It also moved the hard problems. When a network, rather than a warp, decides where every thread goes, fine detail has to survive the model’s compressed latent space and many steps of denoising, and small text and logos are exactly what tends not to.

The data behind it

Progress in the field has been driven by a small number of public datasets of paired images: a garment, and a person wearing it. VITON-HD released one alongside its paper, and Dress Code extended coverage from tops to lower-body garments and dresses [12]. They made the research possible. They are also licensed for non-commercial or academic use only, which matters a great deal for anyone building a product. We come back to this in Where training data comes from.

What is still unsolved

Today’s best systems produce images that look real. Looking real is not the same as being right. What remains hard:

  • Fine detail. Printed text, logos and small patterns are still the first things to break.
  • Fit and drape. How a specific fabric hangs on a specific body is mostly guessed, not reasoned about.
  • Consistency. The same outfit in two poses, or across a video, often disagrees with itself.
  • Speed. Many denoising steps with large networks are expensive to run for every shopper.

These are the problems our research team works on. Each has its own write-up on our open problems page.

References

  1. Han et al. VITON: An Image-based Virtual Try-on Network. CVPR, 2018.
  2. Wang et al. Toward Characteristic-Preserving Image-based Virtual Try-On Network. ECCV, 2018.
  3. Choi et al. VITON-HD: High-Resolution Virtual Try-On via Misalignment-Aware Normalization. CVPR, 2021.
  4. Lee et al. High-Resolution Virtual Try-On with Misalignment and Occlusion-Handled Conditions. ECCV, 2022.
  5. Zhu et al. TryOnDiffusion: A Tale of Two UNets. CVPR, 2023.
  6. Morelli et al. LaDI-VTON: Latent Diffusion Textual-Inversion Enhanced Virtual Try-On. ACM Multimedia, 2023.
  7. Kim et al. StableVITON: Learning Semantic Correspondence with Latent Diffusion Model for Virtual Try-On. CVPR, 2024.
  8. Choi et al. Improving Diffusion Models for Authentic Virtual Try-on in the Wild. ECCV, 2024.
  9. Xu et al. OOTDiffusion: Outfitting Fusion based Latent Diffusion for Controllable Virtual Try-on. AAAI, 2025.
  10. Chong et al. CatVTON: Concatenation Is All You Need for Virtual Try-On with Diffusion Models. ICLR, 2025.
  11. Zhu et al. M&M VTO: Multi-Garment Virtual Try-On and Editing. CVPR, 2024.
  12. Morelli et al. Dress Code: High-Resolution Multi-Category Virtual Try-On. ECCV, 2022.

JOIN THE RESEARCH TEAM

Work on the open problems with us.

Students, researchers and engineers. The application takes about two minutes.