2.9Video virtual try-on
Video try-on has moved in three waves. In 2024, methods such as ViViD added motion layers to Stable Diffusion 1.5 and ran a second network over the garment [78]. In 2025, video diffusion transformers took over: CatV2TON concatenates garment and person in time, MagicTryOn fine-tunes the 14-billion-parameter Wan2.1 model, and DreamVVT tries the garment on a few keyframes first, then generates the video around them [79, 80, 81]. In 2026, almost every new paper fine-tunes a Wan2.1 or Wan2.2 variant, often with LoRA, and the field is moving to mask-free inputs, synthetic training pairs and keyframe-first pipelines [82, 83, 84, 85].
| Method | Year and venue | Base | Key idea | Speed reported | Weights licence |
|---|---|---|---|---|---|
| ViViD [78] | 2024 | SD 1.5 with motion layers | Garment encoder with temporal attention; 9,700-pair dataset | About 200 s per clip | Code Apache-2.0; footage scraped |
| CatV2TON [79] | CVPR 2025 workshop | EasyAnimate DiT | Garment and person concatenated in time; one model for image and video | About 200 s per clip | CC BY-NC-ND |
| MagicTryOn [80] | 2025 | Wan2.1, 14B | Garment tokens with garment-aware position encoding; a Turbo version distilled to 4 steps | 6.7 s per 64 frames (Turbo) | CC BY-NC-SA |
| DreamVVT [81] | 2025, ByteDance | In-house MMDiT | Try on keyframes first, then generate the video from pose and keyframes | 50 steps | Not released |
| BooM-VVT [84] | ACM MM 2026 | Qwen-Image-Edit-2511 keyframes with a Wan-Animate LoRA | Mask-free; picks the keyframes where the garment matters most | About 280 s per 65 frames | CC BY-NC-SA |
| UniVVT [85] | 2026 | Wan2.1 with a vision-language model | No mask, pose or warping at inference | Conditioning in 2 to 3 s | Not released |
| LiveVVT [71] | 2026 | Wan2.1 1.3B student, 14B teacher | Rolling window with a persistent garment memory, 4 steps | 22.4 fps at 512×384 | No code released as of September 2026 |
| Vidu S2-Editing [86] | 2026, closed | In-house | Bidirectional editor adapted to causal streaming | 25 to 42 fps at 720p on several GPUs | Closed |
Speeds were measured on different GPUs, resolutions and clip lengths and are not comparable. Reported VFID scores differ by about a hundredfold between papers, and one widely used VFID toolkit scores only the first 10 frames of each video [84].
The papers name the same unsolved problems. Garment detail degrades in close-ups and at higher resolution, and no video try-on paper reports a metric for printed text or logos over time. Fast motion, occlusion and back views still break models. Long videos drift at window boundaries. Per-frame masks and pose maps cost about a second a frame to compute, which is more than a distilled generator needs, so live try-on has to be mask-free [85, 71].
2.10Real-time video generation
Standard image and video generators start from noise and clean it up over 20 to 50 passes, taking seconds for one image. Live try-on needs a new image every 33 to 66 milliseconds. The field got there with four published ideas, and every leading real-time system uses some mix of them:
- Fewer passes, through distillation. A slow, high-quality teacher trains a fast student to do the same job in 1 to 4 passes [87].
- Frame-by-frame generation with memory. Each frame is drawn from the live camera frame and the model's own recent frames, kept in a key-value cache, so the garment does not flicker or change design between frames [88].
- Training on its own mistakes. Small errors normally pile up frame after frame. Self-forcing and related methods train the model on its own imperfect outputs, so it learns to correct drift instead of amplifying it [89].
- Working small and engineering hard. Models draw a compressed image and then upscale it. They run on top GPUs with custom kernels, 8-bit or 4-bit arithmetic, sparse attention and a direct stream with no queue [90, 91].
| System | Year and origin | How it runs in real time | Reported speed | Weights |
|---|---|---|---|---|
| CausVid [87] | 2025, research | 4-step causal student distilled from a slow bidirectional teacher | About 17 fps on one H100 on Wan2.1 1.3B | Non-commercial |
| Self-Forcing [89] | 2025, research | Trained on its own rolled-out frames with a key-value cache, 4 steps | 17 fps at 832×480 on one H100 | Apache-2.0 |
| LongLive 1.0 [92] | 2025, NVIDIA | Permanent frame sink plus short-window attention, trained on 60-second rollouts | 20.7 fps at 832×480 on one H100 | Code Apache-2.0; weights non-commercial |
| Krea Realtime 14B [93] | 2025, Krea | Self-forcing scaled to Wan2.1 14B, 4 steps | 11 fps on one B200 | Weights Apache-2.0; code non-commercial |
| StreamDiffusionV2 [94] | MLSys 2026 | Training-free serving with a rolling key-value cache and sink tokens | 58 to 64 fps on 4 H100s | Apache-2.0 |
| JoyAI-Video-Edit [95] | 2026, JD | 16B live instruction editor in 2 steps; its demo includes clothing changes | About 30 fps at 720×1248 on one B200 | Apache-2.0 |
| MirageLSD and Lucy 2.5 [88, 90] | 2025 to 2026, Decart | Frame-by-frame generation with history augmentation and custom kernels | 24 to 30 fps, under 40 ms per frame | Closed |
| Seaweed APT2 [91] | 2025, ByteDance | One network evaluation per latent frame after adversarial post-training | 736×416 at 24 fps on one H100 | Closed |
| LiveVVT [71] | 2026, research | Try-on: rolling window with a persistent garment memory, 4 steps | 22.4 fps at 512×384 | No code released |
Speeds are each team's own figures on different hardware. Several closed systems use more than one GPU per stream at 720p.
The public recipe for a small team follows from these results: start from an open video diffusion transformer, train a garment-conditioned video editor, convert it into a causal student with self-forcing and step distillation, and keep a persistent memory of the garment. Before any custom kernel work, published systems of this kind reach about 11 to 24 frames per second at around 480p on one high-end GPU [93, 96, 91]. Many popular real-time repositories cannot be used commercially: CausVid and the LongLive weights are non-commercial, and Krea's code is too, although its weights are Apache-2.0 [87, 92, 93]. The Self-Forcing code and the Wan2.1 models are Apache-2.0 [89, 97].
2.11Live try-on products today
Decart is the clear leader in live try-on. Its Lucy VTON 3.5 model, released in August 2026, takes a live camera stream plus an optional garment photo and text prompt, and returns 720p video over WebRTC, the technology video calls use. Its SDK defines the model at 1280×720 and up to 30 frames per second. Sessions use short-lived client tokens, garments can be switched mid-session, and prompts and garment images are moderated on the server. The list price is USD 0.02 a second, or USD 1.20 a minute, and a faster mode costs double [72, 70, 98]. Decart discloses the principles, not the recipe: causal frame-by-frame generation, training on its own outputs, step distillation, custom GPU kernels and low-precision arithmetic [88, 90].
| Organisation | Offering | Live or offline | Output and price | Method disclosed |
|---|---|---|---|---|
| Decart | Lucy VTON 3.5 API; Anywear store widget | Live camera | 720p; USD 1.20 per minute | Principles only |
| Vidu (ShengShu) | S2-Editing | Live | 720p at 25 to 42 fps on several GPUs | Paper [86] |
| Doppl animated try-on, folded into Search | Offline | Short clips from a photo | Not disclosed | |
| FASHN AI | Image-to-Video API | Offline | 5 to 10 s clips up to 1080p | Not disclosed |
| Luma AI | Ray 3.2 video-to-video wardrobe change | Offline | Up to 20 s, up to 1080p | Not disclosed |
| Runway | Aleph 2.0 video editing | Offline | Up to 30 s at 1080p | Not disclosed |
| Kling | Try-on image, then image-to-video | Offline | Clips from a try-on image | Not disclosed |
For FabricVTON, Decart is both the benchmark to beat and a way to prove shopper demand while its own model is built. It will stay ahead on general live video. FabricVTON needs to win only on clothes: prints and logos that stay exact, fabric that looks like the real product, a much lower cost per minute, and a model it controls.
2.12Video base models, data and their licences
The licence rules for images apply to video. Almost every released video try-on checkpoint is non-commercial, but the Wan video models they build on are licensed Apache-2.0. A company can build on Wan, as long as it trains its own weights on its own data. Wan 2.5 and later have no open weights. Several other open video models carry traps: LTX-2 charges above USD 10 million of revenue and bars competing products, and HunyuanVideo's licence does not apply in the EU, UK or South Korea [99, 100].
| Model | Released | Size | Relevant capability | Licence | Commercial use |
|---|---|---|---|---|---|
| Wan2.1, with VACE and Fun-Control [97] | 2025 | 1.3B and 14B | Reference-to-video, masked video-to-video and pose control; the base of most try-on and real-time work | Apache-2.0 | Yes |
| Wan2.2 and Wan2.2-Animate [101] | 2025 to 2026 | 5B and 14B | 720p at 24 fps; pose-driven character replacement | Apache-2.0 | Yes |
| JoyAI-Video-Edit [95] | 2026 | 16B | Live instruction editing of a camera feed | Apache-2.0 | Yes |
| LTX-2 [99] | 2026 | 19B to 22B | Up to 4K; video-to-video and reference control | LTX-2 Community License | Free below USD 10M revenue; anti-competition clause |
| HunyuanVideo 1.5 [100] | 2025 | 8.3B | Image-to-video; try-on listed as an application | Tencent Hunyuan Community | Not in the EU, UK or South Korea |
| CogVideoX 5B | 2024 | 5B | Image-to-video | CogVideoX licence | Capped at one million visits a month |
| Released video try-on models | 2025 to 2026 | Various | MagicTryOn, CatV2TON, BooM-VVT | CC BY-NC-SA or CC BY-NC-ND | No |
| Dataset | Content | Commercial training |
|---|---|---|
| ViViD [78] | 9,700 garment-video pairs from Net-A-Porter footage | No. Tagged Apache-2.0, but the footage is scraped retail content |
| VVT and TikTok-derived sets | Hundreds of catwalk and dance clips | No. Research use or no licence |
| TripVVT-10K [102] | 10,031 synthetic triplets at 720×1280 | No. CC BY-NC 4.0 |
| MV-Fashion [103] | 80 consented subjects, 754 garments, 68 cameras | No. CC BY-NC-SA 4.0 |
| In-house sets | DreamVVT 69,643 videos; Fashion-VDM 52,000 videos | Not released |
No public video try-on dataset is clearly cleared for commercial training. The approach the field has settled on is the synthetic-triplet idea applied to video: keep a real video as the target, create the input by re-dressing it synthetically, and pair it with the real garment photo [102, 85]. The model then only ever learns to reproduce real footage.
2.13Gap analysis
| Research area | Status | Closest prior work | Implication for this programme |
|---|---|---|---|
| Garments beyond the stitched Western wardrobe: draped, wrapped and regional dress | Open | BD-VITON (1,013 pairs, older models); DIVA | Core of Papers 1 and 2 |
| Fairness by skin tone and body shape | Open | No audit found; slim-body bias noted by SiCo | Built into Paper 1 |
| Garment-faithful live try-on | Open | LiveVVT (512×384, needs masks, no code); closed engines | Core of WP7 and Paper 4 |
| Print and logo stability over time in video | Open | No video try-on paper reports a text or logo metric | Evaluation axis and Paper 3 |
| Commercially clean video try-on data | Open | All public sets are non-commercial or scraped | WP6 data engine |
| Mask-free video try-on | Partly addressed | UniVVT, BooM-VVT | Required for live use |
| Try-off and pseudo-pairs beyond stitched garments | Partly addressed | TryOffDiff [57], Voost, for stitched garments | Method component of Paper 2 |
| Real shoppers' photos | Partly addressed | StreetTryOn; in-the-wild methods such as BooW-VTON | Photo condition as a benchmark axis |
| Size- and fit-aware try-on | Crowded | FIT [58], FitControler [59] | Evaluation axis only |
| Fast, few-step image try-on | Crowded | DirectTryOn, FastFit | Engineering work; optional cost report |
| Layering and multiple garments | Crowded | Layering VTON, Garments2Look, OmniTry | Outerwear and multi-piece outfits as benchmark axes |
| Mask-free image try-on and body preservation | Crowded | BooW-VTON [60], FASHN VTON 1.5 | Personal and cultural details as a benchmark axis |
| General benchmarks and judges | Crowded | OpenVTON-Bench, VTON-QBench, TryOnReward | Calibrate existing judges rather than build new ones |
2.14Summary of the current understanding
Sources in this chapter
- [6]Virtual Try-On for Cultural Clothing: A Benchmarking Study (BD-VITON). arXiv:2603.07291, March 2026. Paper licensed CC BY-NC-SA 4.0; no separate dataset licence stated. arxiv.org/abs/2603.07291
- [15]JD.com. Oxygen-TryOn. arXiv:2607.21694, July 2026. arxiv.org/abs/2607.21694
- [16]Feng, Chen, Shan and Kemelmacher-Shlizerman. Layering Virtual Try-On. ECCV 2026. arXiv:2607.22924. arxiv.org/abs/2607.22924
- [20]DIVA: Indian virtual try-on (IndicViton). ECCV 2024 Workshops, Springer. link.springer.com/chapter/10.1007/978-3-031-91569-7_23
- [21]Awesome Try-On Models (curated list of try-on research), updated 3 September 2026. github.com/Zheng-Chong/Awesome-Try-On-Models
- [22]SiCo: size-controllable virtual try-on. DIS 2025. arXiv:2408.02803. arxiv.org/abs/2408.02803
- [34]OmniTry: mask-free virtual try-on for wearable objects. arXiv:2508.13632. arxiv.org/abs/2508.13632
- [43]DirectTryOn. arXiv:2605.12939, May 2026. arxiv.org/abs/2605.12939
- [52]Garments2Look. CVPR 2026. arXiv:2603.14153. arxiv.org/abs/2603.14153
- [53]OpenVTON-Bench. arXiv:2601.22725, January 2026. arxiv.org/abs/2601.22725
- [54]VTBench: a hierarchical virtual try-on benchmark. arXiv:2505.19571, May 2025. arxiv.org/abs/2505.19571
- [55]VTONQA. arXiv:2601.02945, January 2026. arxiv.org/abs/2601.02945
- [56]VTON-IQA and VTON-QBench. arXiv:2603.13057, March 2026. arxiv.org/abs/2603.13057
- [57]Velioglu et al. TryOffDiff: virtual try-off with diffusion models. BMVC 2025. arXiv:2411.18350. arxiv.org/abs/2411.18350
- [58]Karras et al. FIT: a fit-aware virtual try-on dataset. SIGGRAPH 2026. arXiv:2604.08526. arxiv.org/abs/2604.08526
- [59]FitControler. ECCV 2026. arXiv:2512.24016. arxiv.org/abs/2512.24016
- [60]BooW-VTON: mask-free in-the-wild virtual try-on. CVPR 2025. arXiv:2408.06047. arxiv.org/abs/2408.06047
- [70]Decart. Platform pricing: Lucy VTON realtime at USD 0.02 per second, 2026. docs.platform.decart.ai/getting-started/pricing
- [71]LiveVVT: High-Fidelity Video Virtual Try-On in Real Time. arXiv:2608.26714, August 2026. arxiv.org/abs/2608.26714
- [72]Decart. Realtime virtual try-on (Lucy VTON 3.5) documentation, 2026. docs.platform.decart.ai/models/realtime/virtual-try-on
- [74]Google Labs. Doppl help centre: the app closed on 30 April 2026 and try-on moved into Search. support.google.com/labs/answer/16537062?hl=en
- [75]FASHN AI. Image-to-Video API reference. docs.fashn.ai/api-reference/image-to-video
- [76]Luma AI. Ray 3.2 video-to-video, May 2026. lumalabs.ai/learning-center/articles/ray-3-2-video-to-video
- [77]Runway. Aleph 2.0 video editing. runway.com/product/aleph-2
- [78]Fang et al. ViViD: Video Virtual Try-On using Diffusion Models. arXiv:2405.11794, 2024. arxiv.org/abs/2405.11794
- [79]Chong et al. CatV2TON: temporal concatenation for image and video try-on. CVPR 2025 Workshops. arXiv:2501.11325. arxiv.org/abs/2501.11325
- [80]MagicTryOn: video virtual try-on on Wan2.1. arXiv:2505.21325, 2025. arxiv.org/abs/2505.21325
- [81]DreamVVT: keyframe-first video virtual try-on. ByteDance. arXiv:2508.02807, 2025. arxiv.org/abs/2508.02807
- [82]KeyTailor: instruction-guided keyframes for video try-on. CVPR 2026. arXiv:2512.20340. arxiv.org/abs/2512.20340
- [83]Vanast: animation and garment transfer from one image. CVPR 2026. arXiv:2604.04934. arxiv.org/abs/2604.04934
- [84]BooM-VVT: mask-free video try-on with garment-sensitive keyframes. ACM MM 2026. arXiv:2609.04120. arxiv.org/abs/2609.04120
- [85]UniVVT: video try-on without masks, pose or warping. arXiv:2608.05745, August 2026. arxiv.org/abs/2608.05745
- [86]Vidu S2-Editing: real-time video editing adapted to causal streaming. arXiv:2609.11638, September 2026. arxiv.org/abs/2609.11638
- [87]Yin et al. From Slow Bidirectional to Fast Autoregressive Video Diffusion Models (CausVid). arXiv:2412.07772. arxiv.org/abs/2412.07772
- [88]Decart. MirageLSD: live stream diffusion technical report, July 2025 (archived copy). web.archive.org/web/20250918033439/https://about.decart.ai/publications/mirage
- [89]Huang et al. Self Forcing: Bridging the Train-Test Gap in Autoregressive Video Diffusion. arXiv:2506.08009. arxiv.org/abs/2506.08009
- [90]Decart. Lucy 2.5: raising the bar for live AI, July 2026. decart.ai/publications/lucy-2-5-raising-the-bar-for-live-ai
- [91]ByteDance Seed. Seaweed APT2: autoregressive adversarial post-training for real-time video. seaweed-apt.com/2
- [92]NVIDIA. LongLive: Real-time Interactive Long Video Generation. arXiv:2509.22622. arxiv.org/abs/2509.22622
- [93]Krea. Krea Realtime 14B. www.krea.ai/blog/krea-realtime-14b
- [94]StreamDiffusionV2. MLSys 2026. arXiv:2511.07399. arxiv.org/abs/2511.07399
- [95]JD. JoyAI-Video-Edit: real-time instruction video editing (Apache-2.0). arXiv:2608.03974. arxiv.org/abs/2608.03974
- [96]Pika. PikaStream 1.0: real-time video chat, April 2026. pika.art/blog/introducing-real-time-video-chat
- [97]Wan Team. Wan2.1 open video foundation models (Apache-2.0). github.com/Wan-Video/Wan2.1
- [98]Decart. Network requirements for realtime sessions. docs.platform.decart.ai/integrations/network-requirements
- [99]Lightricks. LTX-2.x Community License. github.com/Lightricks/LTX-2/blob/main/LICENSE-2_x
- [100]Tencent. HunyuanVideo 1.5 licence (Tencent Hunyuan Community License). github.com/Tencent-Hunyuan/HunyuanVideo-1.5/blob/master/LICENSE
- [101]Wan Team. Wan2.2 and Wan2.2-Animate (Apache-2.0). huggingface.co/Wan-AI/Wan2.2-Animate-14B
- [102]TripVVT: synthetic triplets for video try-on, and the TripVVT-10K dataset (CC BY-NC 4.0). arXiv:2604.27958. arxiv.org/abs/2604.27958
- [103]MV-Fashion: multi-view studio fashion video dataset (CC BY-NC-SA 4.0). CVPR 2026. arXiv:2603.08147. arxiv.org/abs/2603.08147
