Skip to content
Clothsy AI Talk to us

Chapter 3AThe programme9 min readVersion 1.2 · 30 September 2026

Technical approach

Measure first, build commercially clean, and use the image model as the teacher for video.

3.1Objectives

Table 16 Programme objectives
IDObjectiveHow success is measuredMonth
O1Build and release a licensed, consented global try-on benchmarkAt least 2,000 image test pairs across at least eight garment families from at least five world regions, labelled for garment attributes, Monk Skin Tone and body shape5
O2Audit current try-on systems on the benchmarkAt least ten open and commercial systems, a human study with raters from several regions, and the agreement between automatic metrics and those raters6
O3Build a commercially clean training pipeline and image modelAt least 20,000 real and synthetic pairs with provenance, and a LoRA fine-tune of Qwen-Image-Edit-25118
O4Beat the best open image baseline at a servable costA statistically significant win in human-rated garment fidelity with no loss on standard metrics, and a distilled model within twice today's 6.5 seconds on an L40S11
O5Build a commercially clean video try-on modelA video training set of at least 7,000 real clips and 20,000 verified synthetic pairs with provenance, and an offline model that beats open video baselines on garment fidelity over time on the 300-clip hold-out13
O6Deliver live try-on on FabricVTON's own modelAt least 15 frames per second at 512p with the first frame in under 2 seconds on one GPU, garment fidelity close to the offline model, and a cost per stream-minute under a tenth of the rented engine16

3.2Design principles

  • Measure before modelling. The benchmark and evaluation harness come first, so every model decision is scored across garment families, photo conditions and diverse bodies.
  • Commercially clean by construction. Only Apache-2.0 or MIT bases and teachers, licensed or consented data, and a provenance record for every training sample.
  • Coverage and fairness, not raw scale. A small team cannot match 50 million images. It can build the most diverse and best-measured data for the garments, bodies and photos that large players' published results leave out.
  • Large teacher, small student. Train the best model on a 20-billion-parameter base, then distil it into a model cheap enough to serve on L4 and L40S GPUs.
  • Open science where it helps the field. Publish methods, the benchmark and the evaluation toolkit. Keep the production model and licensed training data proprietary.
  • Image first, live video as the goal. The image model comes first because it teaches the video model: it re-dresses real clips to make training pairs. Every image-track decision is judged by whether it also serves the video track.

3.3System overview

Figure 1 shows how the seven work packages connect. The image track turns licensed data into a servable image model. The video and live track reuses that model as a teacher and ends in live try-on. The measurement track scores every checkpoint in both.

Image track

Data sourcesLicensed photoshootsLicensed cataloguesApache-2.0 teachersTry-off images
WP2 Data enginePairs and tripletsQuality filtersProvenance ledgerConsent records
WP3 Image modelQwen-Image-Edit-2511LoRA, rank 32Multi-view garmentsWear instruction
WP4 Tune and distilGlobal rater panelReward or judge4B student modelSpeed lab checks
Photo servingProduction serviceL4 and L40S GPUsCost per try-onmeasured

The image model re-dresses video frames as a teacher, and the same distillation recipe carries over.

Video and live track

Video sourcesConsented shootsLicensed motionPhone clipsWebcam clips
Video data engineReal target clipsRe-dressed inputsTemporal filtersHuman QA
WP6 Video modelOpen video baseGarment memoryPerson video inOffline quality first
WP7 Live modelCausal, frame by frameSelf-forcing training1 to 4 stepsKV cache
Live servingWebRTC streamOwn GPU servers15 to 20 fps targetCost per minute

The evaluation harness scores every checkpoint in both tracks.

Measurement and research-output track

WP1 Benchmark and audit2,000+ image test pairs300-clip video hold-out8+ garment families, 5+ regionsSkin tone and body strata
Evaluation harnessImage and video metricsHuman ratingsLive fps, latency, driftFairness measures
WP5 OutputsPapers 1 to 4Open toolkitBenchmark releaseOwn live model
Figure 1Proposed system overview: the image track, the video and live track, and the measurement track.
Table 17 Work packages
WPNamePurposeLead role
WP1Benchmark and fairness auditMeasure how current systems perform across garment families, regions, photo and video conditions and diverse bodiesEvaluation lead
WP2Data engineBuild commercially clean image training pairs with provenanceData and consent lead
WP3Image modelFine-tune Qwen-Image-Edit-2511 for try-on of any garmentModelling lead
WP4Tuning and distillationAlign with human preference and make image serving affordableModelling and infrastructure leads
WP5Dissemination and transferPapers, releases and product integrationFounders
WP6Video try-on modelBuild commercially clean video data and an offline video try-on modelVideo lead
WP7Live try-onTurn the video model into a live, streaming model and serve itVideo and infrastructure leads

3.4WP1: Benchmark and fairness audit

The benchmark is a test set built from new, licensed photography and licensed partner catalogues. Each item pairs a photograph of a consenting model with a garment presented in one of four ways: flat or unfolded, folded as sold, on a hanger or mannequin, or worn by another person, with a back view wherever it differs from the front. This lets the benchmark measure how much the garment presentation drives failure, separately from the model. Person photographs come in two conditions, studio and real-shopper (phone selfies, mirror shots and dim light), so the benchmark also measures how much the photo drives failure. Subjects are balanced across all ten groups of the Monk Skin Tone scale and across body shapes, including plus sizes [24].

12345678910
Figure 3.1The ten groups of the Monk Skin Tone scale, created by Dr Ellis Monk with Google. Benchmark subjects are balanced across all ten, and across body shapes including plus sizes.
Table 18 Benchmark garment taxonomy and labelled attributes
Garment familyExamplesAttributes labelled
TopsT-shirts, shirts, blouses, knitwear, hoodiesFit, neckline, sleeve length, print, logo and text fidelity
BottomsJeans, trousers, skirts, shortsRise, length, silhouette, waistband, wash and texture
One-piecesDresses, jumpsuitsLength, flare, silhouette, pattern fidelity
Outerwear and layeringCoats, jackets, blazers and cardigans, worn open or closed over an existing outfitLayer order, inner-layer preservation, closures, collar and lapels
Tailoring and formalwearSuits, waistcoats, formal gowns, sherwaniStructure, closures, fabric drape, embroidery
Multi-piece outfitsCo-ords, kurta sets, lehenga with choli and dupattaConsistency across pieces, layering, scarf or dupatta placement
Draped and wrapped garmentsSaree in several drape styles, dhoti, sarong, shawls, wrap dressesDrape and fold, pleats, border continuity, decorated ends such as the pallu
Regional dressKimono and yukata, hanbok, qipao, kaftan and abaya, dashiki, salwar kameezStructure, closures, embroidery and motif fidelity, cultural details
Preservation checksAll itemsFace, hands, hair, tattoos, jewellery, headscarves and turbans, mehndi and bindi, skin tone, body shape

At least ten systems will be audited without retraining. Open specialists include FASHN VTON 1.5, CatVTON, IDM-VTON, Leffa, FitDiT and OmniTry. Open general editors include Qwen-Image-Edit-2511 and FLUX.2 klein base 4B. Commercial APIs, such as FASHN's API and Google's virtual-try-on-001, will be included where their terms allow evaluation. Checkpoints with non-commercial licences will be run only where their terms allow research evaluation, and their results will appear only in publications.

Every subject signs a consent form covering AI training, evaluation and publication, in line with India's Digital Personal Data Protection Act and, for subjects in the United Kingdom or the European Union, the UK GDPR and the GDPR. Faces are anonymised in the public release where consent does not cover publication, and the release carries a datasheet.

The benchmark also has a video hold-out of about 300 clips, with people and garments kept out of all training data: 100 studio clips, 100 consented phone clips, and 100 live-camera stress clips recorded on a webcam, with low light, fast spins, hands and bags across the body, back turns and sequences of a minute or more.

3.5WP2: Data engine

Table 19 Training data sources and licence position
SourceWhat it providesLicence positionRole
Licensed photoshootsProduct views (front, back and detail) plus two to four on-model poses per garment, on several models; folded, hanger and close-up views where a garment is sold that wayOwned, with model releases that cover AI trainingReal pairs for high-value fine-tuning
Brand catalogue licencesOn-model and product photos from direct-to-consumer partners in several regionsLicensed with an explicit machine-learning clauseVolume and diversity
Licensed stock photographyPerson photographs across regions, skin tones and body shapesLicensed, with model releases that cover AI trainingSubject diversity
Synthetic tripletsA real photo re-dressed by a teacher model becomes the input; the real photo stays the targetTeachers are Apache-2.0: FASHN VTON 1.5 and Qwen-Image-Edit-2511Mask-free training at scale
Try-off generationA flat garment image generated from a worn photoGenerated with an Apache-2.0 modelPairs from catalogues that only have on-model shots
Public research datasetsVITON-HD, DressCodeNon-commercialEvaluation only, never training

Every sample enters a provenance ledger that records its source, licence, consent reference and any teacher model. Filters remove low-quality and duplicate pairs, and each batch is spot-checked by hand. The target is at least 20,000 training pairs, balanced across garment families and subjects. Generating 500,000 candidate synthetic triplets with the current teacher would cost about USD 2,000, based on the measured 6.5 seconds per try-on on an L40S billed at USD 2.27 per hour [7].

3.6WP3: Model adaptation

Base model

Qwen-Image-Edit-2511 is the base model for four reasons. It is licensed Apache-2.0, with no restriction on outputs. It accepts several input images natively, so the person and each garment view can be supplied separately. It is the base used by strong 2026 try-on work, including Layering VTON at ECCV 2026. And it is supported by the main open training tools, including musubi-tuner [48, 16, 61]. FLUX.2 klein base 4B, also Apache-2.0, is kept as the student model for serving.

Conditioning on several views of one garment

The central modelling idea addresses the ill-posed inputs described in the problem chapter. Instead of one garment photograph, the model receives several views of the same garment: the front and the back, a close-up of any print, logo or embroidery, and, for a draped or wrapped garment, the unfolded fabric and its decorated border or end. For a multi-piece outfit it receives each piece. A short instruction names the garment type and how it is worn, for example a coat worn open over a shirt, or a saree in a Nivi drape. The research question is whether this multi-view conditioning, together with the wear instruction, recovers detail that a single product photograph cannot supply.

Training recipe

Table 20 Initial training configuration
SettingValueBasis
AdapterLoRA, rank 32 and alpha 32, on the query, key and value projectionsLayering VTON
OptimiserLearning rate 1×10⁻⁴, constant, no warm-upLayering VTON
Batch and precisionGlobal batch 32, bf16 mixed precisionLayering VTON
Stage 1About 20,000 steps on synthetic tripletsLayering VTON scale
Stage 2About 5,000 steps on licensed real pairsLayering VTON scale
Stage 3High-resolution stage for fine print, logos and embroideryTAMF-VTON, FitDiT
HardwareOne H200, or one H100 with FP8 weightsLayering VTON; musubi-tuner reports about 42 GB without options and 30 GB with FP8
AblationsSingle versus multi-view garments; with and without the wear instruction; frequency-domain texture loss; attention-correspondence lossFitDiT, TAMF-VTON, Leffa

3.7WP4: Preference tuning and distillation

Once supervised fine-tuning works, a panel of raters from several regions will compare outputs in pairs. Their judgements will train a small reward model or calibrate a vision-language judge, which then guides reinforcement learning of the kind used by Tstars-Tryon and Oxygen-TryOn [3, 15]. This stage improves the product and is optional for the papers.

A 20-billion-parameter model is expensive to serve. The programme will first test the Apache-2.0 four-step Lightning LoRA for Qwen-Image-Edit-2511, then distil the fine-tuned model into FLUX.2 klein base 4B [62]. The target is a distilled model within twice today's measured 6.5 seconds on an L40S, measured with FabricVTON's existing speed laboratory. The same work package replaces a third-party component in the current serving pipeline with a commercially licensed one.

Sources in this chapter

  1. [3]Alibaba Taobao. Tstars-Tryon 1.0. arXiv:2604.19748, April 2026. arxiv.org/abs/2604.19748
  2. [7]FabricVTON. Internal technical documentation: What Is Built (updated 18 and 25 September 2026) and Project Status (24 September 2026). Available to reviewers on request.
  3. [15]JD.com. Oxygen-TryOn. arXiv:2607.21694, July 2026. arxiv.org/abs/2607.21694
  4. [16]Feng, Chen, Shan and Kemelmacher-Shlizerman. Layering Virtual Try-On. ECCV 2026. arXiv:2607.22924. arxiv.org/abs/2607.22924
  5. [24]Monk, E. and Google. The Monk Skin Tone Scale. skintone.google
  6. [27]Jiang et al. FitDiT: Advancing the Authentic Garment Details for High-Fidelity Virtual Try-On. arXiv:2411.10499. arxiv.org/abs/2411.10499
  7. [36]TAMF-VTON. arXiv:2607.14807, July 2026. arxiv.org/abs/2607.14807
  8. [48]Qwen Team. Qwen-Image-Edit-2511 model card (Apache-2.0). huggingface.co/Qwen/Qwen-Image-Edit-2511
  9. [61]kohya-ss. musubi-tuner documentation for Qwen-Image training. github.com/kohya-ss/musubi-tuner/blob/main/docs/qwen_image.md
  10. [62]LightX2V. Qwen-Image-Edit-2511-Lightning (few-step LoRA). huggingface.co/lightx2v/Qwen-Image-Edit-2511-Lightning