Before a model can dress someone, it has to understand the photo it was given.
Every try-on starts with perception. The model needs to know where the person is, how they are standing, which pixels are skin, hair, background and clothing, and roughly what shape the body is underneath. It needs to understand the product photo too: what kind of garment it is, where the collar, sleeves and hem are, and which details matter.
These are classic computer-vision problems: human parsing, pose estimation, segmentation and dense body correspondence. Try-on stresses them in unusual ways. Loose, layered or dark clothing hides the body. Product photos arrive in every style, from flat lays and ghost mannequins to model shots. And errors compound: a parsing mistake at the start becomes a visible artifact at the end.
Much of this layer is built on strong open research, such as pose estimators, human parsers and promptable segmentation models. Our work is in making it reliable on real shopper photos and real catalogues, and in knowing when an input will not produce a good result.



Keep from the personFace and identity, pose, body shape, skin, hair, background.
Take from the productGarment shape, colour, fabric texture, print, logos, details like buttons and seams.
Questions we are exploring
- How do we estimate body shape reliably under loose or layered clothing?
- Can one model understand garment structure (collar, sleeves, closures, hem) from any style of product photo?
- Can we tell, before generating anything, that a photo will give a poor result, and tell the person why?
Further reading
- Cao et al. OpenPose: Realtime Multi-Person 2D Pose Estimation Using Part Affinity Fields. IEEE TPAMI, 2021.
- Güler, Neverova & Kokkinos DensePose: Dense Human Pose Estimation In The Wild. CVPR, 2018.
- Yang et al. Effective Whole-body Pose Estimation with Two-stages Distillation. ICCV Workshops (CV4Metaverse), 2023.
- Li et al. Self-Correction for Human Parsing. IEEE TPAMI, 2022.
- Ravi et al. SAM 2: Segment Anything in Images and Videos. ICLR, 2025.


