How do you know whether a try-on is good? The honest answer is that the standard numbers only tell part of the story. A result can score well and still show the wrong logo, a pattern at the wrong scale, or a colour a shade off. Those are the errors a shopper notices first, and they are the ones that matter when the image is standing in for the real product. This post walks through what the common metrics actually measure, where they fall short for texture, and how we think a fidelity benchmark should be built.
Paired and unpaired
Try-on papers usually report results in two settings.
- Paired. The model re-dresses a person in the garment they are already wearing. Because the original photo is the ground truth, you can compare the output to it pixel by pixel. Test sets such as those released with VITON-HD and Dress Code are built this way [1][2].
- Unpaired. The model puts a different garment on the person. This is the real use case, and there is no ground-truth photo to compare against, so only whole-set statistics apply.
| Pixels and structure | Perceptual features | Whole-set statistics | People | |
|---|---|---|---|---|
| PairedSame garment, ground truth exists | SSIMLocal luminance, contrast and structure. | LPIPSDISTSDeep-feature distance, calibrated to human judgement. | Not needed when a reference exists. | Side-by-sideWhich result is closer to the product? |
| UnpairedNew garment, no ground truth | No reference to compare against. | CLIP-ISemantic similarity to the product image. | FIDKIDHow realistic the set looks overall. | PreferenceRealism and faithfulness, judged by people. |
What each metric sees
SSIM
The structural similarity index compares local windows of two images on luminance, contrast and structure [3]. It is interpretable and sensitive to blur and misalignment. It also penalizes a result that is shifted by a few pixels as harshly as one that is wrong, and it has no notion of which regions matter.
LPIPS
LPIPS measures distance between deep network features and was calibrated against human judgements of image similarity [4]. It tracks perceived differences better than SSIM, which is why it became standard for paired try-on evaluation.
DISTS
DISTS was designed to unify structure and texture similarity, with explicit tolerance to texture resampling [5]. Two photos of the same grass, sampled differently, score as similar. That is the right behaviour for natural textures. For a garment, some of that tolerance is a risk: a printed motif that has moved or changed can still look like the same texture.
FID and KID
The Fréchet Inception Distance compares the statistics of deep features over a whole set of real images and a whole set of generated ones [6]. The Kernel Inception Distance does the same with an unbiased estimator that is better behaved on smaller sets [7]. Both measure whether the generated images, as a population, look like real photos. Neither says anything about whether a particular output shows the right garment.
CLIP similarity
Embedding the product image and the try-on with CLIP and comparing them gives a semantic similarity score that works in the unpaired setting [8]. It is good at “is this the same kind of garment?” and weak at “is this the same print?”
What they miss
For texture fidelity specifically, the gaps line up:
- Area. The garment is part of the image, and printed detail is a small part of the garment. Whole-image scores average the error away.
- Meaning. A logo with one wrong letter is a large error to a person and a tiny one to a pixel or feature distance.
- Sets versus samples. FID and KID can improve while individual results get less faithful, because they reward realism across the set, not correctness per image.
- Easy pairs. The paired setting keeps the person’s pose and the garment’s placement, which is the easiest case. The hardest cases, new garments on new poses, are the ones with no reference.
A framework for fidelity
We think a fidelity benchmark should fail a model for the things a shopper would reject. In practice that means layering several measurements rather than trusting any single one:
- Score the garment, not the image. Compute paired metrics such as LPIPS and DISTS on the garment region only, and separately on its printed and detailed areas.
- Check what must be exact. Text and logos should still read correctly, and colour differences should be measured in a perceptual colour space, not by eye.
- Keep identity separate. A try-on that changes the person’s face is a different failure from one that changes the shirt, and should be measured on its own.
- Ask people. Side-by-side preference studies with a fixed protocol are the ground truth the other numbers are trying to approximate. Automated judges, including vision-language models, are useful for regression testing once they have been checked against those human ratings.
- Keep the set-level scores as a sanity check. FID and KID still catch a model whose outputs stop looking like photographs at all.
None of this is exotic. It is the discipline of measuring the thing you care about, on data the model has not seen, before trying to improve it. It is also one of the principles we work by.
References
- Choi et al. VITON-HD: High-Resolution Virtual Try-On via Misalignment-Aware Normalization. CVPR, 2021.
- Morelli et al. Dress Code: High-Resolution Multi-Category Virtual Try-On. ECCV, 2022.
- Wang, Bovik, Sheikh & Simoncelli Image Quality Assessment: From Error Visibility to Structural Similarity. IEEE Transactions on Image Processing, 2004.
- Zhang et al. The Unreasonable Effectiveness of Deep Features as a Perceptual Metric. CVPR, 2018.
- Ding et al. Image Quality Assessment: Unifying Structure and Texture Similarity. IEEE TPAMI, 2022.
- Heusel et al. GANs Trained by a Two Time-Scale Update Rule Converge to a Local Nash Equilibrium. NeurIPS, 2017.
- Bińkowski et al. Demystifying MMD GANs. ICLR, 2018.
- Radford et al. Learning Transferable Visual Models From Natural Language Supervision. ICML, 2021.
