Skip to content
Clothsy AI Talk to us

Chapter 2BState of the art10 min readVersion 1.2 · 30 September 2026

Video and live try-on

Live try-on has just become possible. How the field got there, who ships it, and what is still open.

2.9Video virtual try-on

Video try-on has moved in three waves. In 2024, methods such as ViViD added motion layers to Stable Diffusion 1.5 and ran a second network over the garment [78]. In 2025, video diffusion transformers took over: CatV2TON concatenates garment and person in time, MagicTryOn fine-tunes the 14-billion-parameter Wan2.1 model, and DreamVVT tries the garment on a few keyframes first, then generates the video around them [79, 80, 81]. In 2026, almost every new paper fine-tunes a Wan2.1 or Wan2.2 variant, often with LoRA, and the field is moving to mask-free inputs, synthetic training pairs and keyframe-first pipelines [82, 83, 84, 85].

Table 10 Selected video try-on methods
MethodYear and venueBaseKey ideaSpeed reportedWeights licence
ViViD [78]2024SD 1.5 with motion layersGarment encoder with temporal attention; 9,700-pair datasetAbout 200 s per clipCode Apache-2.0; footage scraped
CatV2TON [79]CVPR 2025 workshopEasyAnimate DiTGarment and person concatenated in time; one model for image and videoAbout 200 s per clipCC BY-NC-ND
MagicTryOn [80]2025Wan2.1, 14BGarment tokens with garment-aware position encoding; a Turbo version distilled to 4 steps6.7 s per 64 frames (Turbo)CC BY-NC-SA
DreamVVT [81]2025, ByteDanceIn-house MMDiTTry on keyframes first, then generate the video from pose and keyframes50 stepsNot released
BooM-VVT [84]ACM MM 2026Qwen-Image-Edit-2511 keyframes with a Wan-Animate LoRAMask-free; picks the keyframes where the garment matters mostAbout 280 s per 65 framesCC BY-NC-SA
UniVVT [85]2026Wan2.1 with a vision-language modelNo mask, pose or warping at inferenceConditioning in 2 to 3 sNot released
LiveVVT [71]2026Wan2.1 1.3B student, 14B teacherRolling window with a persistent garment memory, 4 steps22.4 fps at 512×384No code released as of September 2026
Vidu S2-Editing [86]2026, closedIn-houseBidirectional editor adapted to causal streaming25 to 42 fps at 720p on several GPUsClosed

Speeds were measured on different GPUs, resolutions and clip lengths and are not comparable. Reported VFID scores differ by about a hundredfold between papers, and one widely used VFID toolkit scores only the first 10 frames of each video [84].

The papers name the same unsolved problems. Garment detail degrades in close-ups and at higher resolution, and no video try-on paper reports a metric for printed text or logos over time. Fast motion, occlusion and back views still break models. Long videos drift at window boundaries. Per-frame masks and pose maps cost about a second a frame to compute, which is more than a distilled generator needs, so live try-on has to be mask-free [85, 71].

2.10Real-time video generation

Standard image and video generators start from noise and clean it up over 20 to 50 passes, taking seconds for one image. Live try-on needs a new image every 33 to 66 milliseconds. The field got there with four published ideas, and every leading real-time system uses some mix of them:

Fewer passesA slow teacher trains a fast student to do in 1 to 4 passes what took 20 to 50.
Memory of recent framesEach frame is drawn from the live camera frame and the model's own recent frames, kept in a key-value cache.
Training on its own mistakesSelf-forcing trains on the model's own imperfect outputs, so small errors are corrected instead of piling up.
Work small, engineer hardDraw a compressed image, upscale it, and run on top GPUs with custom kernels and low-precision arithmetic.
Figure 2.2The four published ideas that took video generation from seconds per image to a new frame every 33 to 66 milliseconds.
  • Fewer passes, through distillation. A slow, high-quality teacher trains a fast student to do the same job in 1 to 4 passes [87].
  • Frame-by-frame generation with memory. Each frame is drawn from the live camera frame and the model's own recent frames, kept in a key-value cache, so the garment does not flicker or change design between frames [88].
  • Training on its own mistakes. Small errors normally pile up frame after frame. Self-forcing and related methods train the model on its own imperfect outputs, so it learns to correct drift instead of amplifying it [89].
  • Working small and engineering hard. Models draw a compressed image and then upscale it. They run on top GPUs with custom kernels, 8-bit or 4-bit arithmetic, sparse attention and a direct stream with no queue [90, 91].
Table 11 Real-time video generation systems
SystemYear and originHow it runs in real timeReported speedWeights
CausVid [87]2025, research4-step causal student distilled from a slow bidirectional teacherAbout 17 fps on one H100 on Wan2.1 1.3BNon-commercial
Self-Forcing [89]2025, researchTrained on its own rolled-out frames with a key-value cache, 4 steps17 fps at 832×480 on one H100Apache-2.0
LongLive 1.0 [92]2025, NVIDIAPermanent frame sink plus short-window attention, trained on 60-second rollouts20.7 fps at 832×480 on one H100Code Apache-2.0; weights non-commercial
Krea Realtime 14B [93]2025, KreaSelf-forcing scaled to Wan2.1 14B, 4 steps11 fps on one B200Weights Apache-2.0; code non-commercial
StreamDiffusionV2 [94]MLSys 2026Training-free serving with a rolling key-value cache and sink tokens58 to 64 fps on 4 H100sApache-2.0
JoyAI-Video-Edit [95]2026, JD16B live instruction editor in 2 steps; its demo includes clothing changesAbout 30 fps at 720×1248 on one B200Apache-2.0
MirageLSD and Lucy 2.5 [88, 90]2025 to 2026, DecartFrame-by-frame generation with history augmentation and custom kernels24 to 30 fps, under 40 ms per frameClosed
Seaweed APT2 [91]2025, ByteDanceOne network evaluation per latent frame after adversarial post-training736×416 at 24 fps on one H100Closed
LiveVVT [71]2026, researchTry-on: rolling window with a persistent garment memory, 4 steps22.4 fps at 512×384No code released

Speeds are each team's own figures on different hardware. Several closed systems use more than one GPU per stream at 720p.

The public recipe for a small team follows from these results: start from an open video diffusion transformer, train a garment-conditioned video editor, convert it into a causal student with self-forcing and step distillation, and keep a persistent memory of the garment. Before any custom kernel work, published systems of this kind reach about 11 to 24 frames per second at around 480p on one high-end GPU [93, 96, 91]. Many popular real-time repositories cannot be used commercially: CausVid and the LongLive weights are non-commercial, and Krea's code is too, although its weights are Apache-2.0 [87, 92, 93]. The Self-Forcing code and the Wan2.1 models are Apache-2.0 [89, 97].

2.11Live try-on products today

Decart is the clear leader in live try-on. Its Lucy VTON 3.5 model, released in August 2026, takes a live camera stream plus an optional garment photo and text prompt, and returns 720p video over WebRTC, the technology video calls use. Its SDK defines the model at 1280×720 and up to 30 frames per second. Sessions use short-lived client tokens, garments can be switched mid-session, and prompts and garment images are moderated on the server. The list price is USD 0.02 a second, or USD 1.20 a minute, and a faster mode costs double [72, 70, 98]. Decart discloses the principles, not the recipe: causal frame-by-frame generation, training on its own outputs, step distillation, custom GPU kernels and low-precision arithmetic [88, 90].

Table 12 Commercial video and live try-on, September 2026
OrganisationOfferingLive or offlineOutput and priceMethod disclosed
DecartLucy VTON 3.5 API; Anywear store widgetLive camera720p; USD 1.20 per minutePrinciples only
Vidu (ShengShu)S2-EditingLive720p at 25 to 42 fps on several GPUsPaper [86]
GoogleDoppl animated try-on, folded into SearchOfflineShort clips from a photoNot disclosed
FASHN AIImage-to-Video APIOffline5 to 10 s clips up to 1080pNot disclosed
Luma AIRay 3.2 video-to-video wardrobe changeOfflineUp to 20 s, up to 1080pNot disclosed
RunwayAleph 2.0 video editingOfflineUp to 30 s at 1080pNot disclosed
KlingTry-on image, then image-to-videoOfflineClips from a try-on imageNot disclosed

For FabricVTON, Decart is both the benchmark to beat and a way to prove shopper demand while its own model is built. It will stay ahead on general live video. FabricVTON needs to win only on clothes: prints and logos that stay exact, fabric that looks like the real product, a much lower cost per minute, and a model it controls.

2.12Video base models, data and their licences

The licence rules for images apply to video. Almost every released video try-on checkpoint is non-commercial, but the Wan video models they build on are licensed Apache-2.0. A company can build on Wan, as long as it trains its own weights on its own data. Wan 2.5 and later have no open weights. Several other open video models carry traps: LTX-2 charges above USD 10 million of revenue and bars competing products, and HunyuanVideo's licence does not apply in the EU, UK or South Korea [99, 100].

Table 13 Open video models relevant to a commercial try-on model
ModelReleasedSizeRelevant capabilityLicenceCommercial use
Wan2.1, with VACE and Fun-Control [97]20251.3B and 14BReference-to-video, masked video-to-video and pose control; the base of most try-on and real-time workApache-2.0Yes
Wan2.2 and Wan2.2-Animate [101]2025 to 20265B and 14B720p at 24 fps; pose-driven character replacementApache-2.0Yes
JoyAI-Video-Edit [95]202616BLive instruction editing of a camera feedApache-2.0Yes
LTX-2 [99]202619B to 22BUp to 4K; video-to-video and reference controlLTX-2 Community LicenseFree below USD 10M revenue; anti-competition clause
HunyuanVideo 1.5 [100]20258.3BImage-to-video; try-on listed as an applicationTencent Hunyuan CommunityNot in the EU, UK or South Korea
CogVideoX 5B20245BImage-to-videoCogVideoX licenceCapped at one million visits a month
Released video try-on models2025 to 2026VariousMagicTryOn, CatV2TON, BooM-VVTCC BY-NC-SA or CC BY-NC-NDNo
Table 14 Video try-on datasets and whether they permit commercial training
DatasetContentCommercial training
ViViD [78]9,700 garment-video pairs from Net-A-Porter footageNo. Tagged Apache-2.0, but the footage is scraped retail content
VVT and TikTok-derived setsHundreds of catwalk and dance clipsNo. Research use or no licence
TripVVT-10K [102]10,031 synthetic triplets at 720×1280No. CC BY-NC 4.0
MV-Fashion [103]80 consented subjects, 754 garments, 68 camerasNo. CC BY-NC-SA 4.0
In-house setsDreamVVT 69,643 videos; Fashion-VDM 52,000 videosNot released

No public video try-on dataset is clearly cleared for commercial training. The approach the field has settled on is the synthetic-triplet idea applied to video: keep a real video as the target, create the input by re-dressing it synthetically, and pair it with the real garment photo [102, 85]. The model then only ever learns to reproduce real footage.

2.13Gap analysis

Table 15 Research gaps in virtual try-on, September 2026
Research areaStatusClosest prior workImplication for this programme
Garments beyond the stitched Western wardrobe: draped, wrapped and regional dressOpenBD-VITON (1,013 pairs, older models); DIVACore of Papers 1 and 2
Fairness by skin tone and body shapeOpenNo audit found; slim-body bias noted by SiCoBuilt into Paper 1
Garment-faithful live try-onOpenLiveVVT (512×384, needs masks, no code); closed enginesCore of WP7 and Paper 4
Print and logo stability over time in videoOpenNo video try-on paper reports a text or logo metricEvaluation axis and Paper 3
Commercially clean video try-on dataOpenAll public sets are non-commercial or scrapedWP6 data engine
Mask-free video try-onPartly addressedUniVVT, BooM-VVTRequired for live use
Try-off and pseudo-pairs beyond stitched garmentsPartly addressedTryOffDiff [57], Voost, for stitched garmentsMethod component of Paper 2
Real shoppers' photosPartly addressedStreetTryOn; in-the-wild methods such as BooW-VTONPhoto condition as a benchmark axis
Size- and fit-aware try-onCrowdedFIT [58], FitControler [59]Evaluation axis only
Fast, few-step image try-onCrowdedDirectTryOn, FastFitEngineering work; optional cost report
Layering and multiple garmentsCrowdedLayering VTON, Garments2Look, OmniTryOuterwear and multi-piece outfits as benchmark axes
Mask-free image try-on and body preservationCrowdedBooW-VTON [60], FASHN VTON 1.5Personal and cultural details as a benchmark axis
General benchmarks and judgesCrowdedOpenVTON-Bench, VTON-QBench, TryOnRewardCalibrate existing judges rather than build new ones

2.14Summary of the current understanding

Sources in this chapter

  1. [6]Virtual Try-On for Cultural Clothing: A Benchmarking Study (BD-VITON). arXiv:2603.07291, March 2026. Paper licensed CC BY-NC-SA 4.0; no separate dataset licence stated. arxiv.org/abs/2603.07291
  2. [15]JD.com. Oxygen-TryOn. arXiv:2607.21694, July 2026. arxiv.org/abs/2607.21694
  3. [16]Feng, Chen, Shan and Kemelmacher-Shlizerman. Layering Virtual Try-On. ECCV 2026. arXiv:2607.22924. arxiv.org/abs/2607.22924
  4. [20]DIVA: Indian virtual try-on (IndicViton). ECCV 2024 Workshops, Springer. link.springer.com/chapter/10.1007/978-3-031-91569-7_23
  5. [21]Awesome Try-On Models (curated list of try-on research), updated 3 September 2026. github.com/Zheng-Chong/Awesome-Try-On-Models
  6. [22]SiCo: size-controllable virtual try-on. DIS 2025. arXiv:2408.02803. arxiv.org/abs/2408.02803
  7. [34]OmniTry: mask-free virtual try-on for wearable objects. arXiv:2508.13632. arxiv.org/abs/2508.13632
  8. [43]DirectTryOn. arXiv:2605.12939, May 2026. arxiv.org/abs/2605.12939
  9. [52]Garments2Look. CVPR 2026. arXiv:2603.14153. arxiv.org/abs/2603.14153
  10. [53]OpenVTON-Bench. arXiv:2601.22725, January 2026. arxiv.org/abs/2601.22725
  11. [54]VTBench: a hierarchical virtual try-on benchmark. arXiv:2505.19571, May 2025. arxiv.org/abs/2505.19571
  12. [55]VTONQA. arXiv:2601.02945, January 2026. arxiv.org/abs/2601.02945
  13. [56]VTON-IQA and VTON-QBench. arXiv:2603.13057, March 2026. arxiv.org/abs/2603.13057
  14. [57]Velioglu et al. TryOffDiff: virtual try-off with diffusion models. BMVC 2025. arXiv:2411.18350. arxiv.org/abs/2411.18350
  15. [58]Karras et al. FIT: a fit-aware virtual try-on dataset. SIGGRAPH 2026. arXiv:2604.08526. arxiv.org/abs/2604.08526
  16. [59]FitControler. ECCV 2026. arXiv:2512.24016. arxiv.org/abs/2512.24016
  17. [60]BooW-VTON: mask-free in-the-wild virtual try-on. CVPR 2025. arXiv:2408.06047. arxiv.org/abs/2408.06047
  18. [70]Decart. Platform pricing: Lucy VTON realtime at USD 0.02 per second, 2026. docs.platform.decart.ai/getting-started/pricing
  19. [71]LiveVVT: High-Fidelity Video Virtual Try-On in Real Time. arXiv:2608.26714, August 2026. arxiv.org/abs/2608.26714
  20. [72]Decart. Realtime virtual try-on (Lucy VTON 3.5) documentation, 2026. docs.platform.decart.ai/models/realtime/virtual-try-on
  21. [74]Google Labs. Doppl help centre: the app closed on 30 April 2026 and try-on moved into Search. support.google.com/labs/answer/16537062?hl=en
  22. [75]FASHN AI. Image-to-Video API reference. docs.fashn.ai/api-reference/image-to-video
  23. [76]Luma AI. Ray 3.2 video-to-video, May 2026. lumalabs.ai/learning-center/articles/ray-3-2-video-to-video
  24. [77]Runway. Aleph 2.0 video editing. runway.com/product/aleph-2
  25. [78]Fang et al. ViViD: Video Virtual Try-On using Diffusion Models. arXiv:2405.11794, 2024. arxiv.org/abs/2405.11794
  26. [79]Chong et al. CatV2TON: temporal concatenation for image and video try-on. CVPR 2025 Workshops. arXiv:2501.11325. arxiv.org/abs/2501.11325
  27. [80]MagicTryOn: video virtual try-on on Wan2.1. arXiv:2505.21325, 2025. arxiv.org/abs/2505.21325
  28. [81]DreamVVT: keyframe-first video virtual try-on. ByteDance. arXiv:2508.02807, 2025. arxiv.org/abs/2508.02807
  29. [82]KeyTailor: instruction-guided keyframes for video try-on. CVPR 2026. arXiv:2512.20340. arxiv.org/abs/2512.20340
  30. [83]Vanast: animation and garment transfer from one image. CVPR 2026. arXiv:2604.04934. arxiv.org/abs/2604.04934
  31. [84]BooM-VVT: mask-free video try-on with garment-sensitive keyframes. ACM MM 2026. arXiv:2609.04120. arxiv.org/abs/2609.04120
  32. [85]UniVVT: video try-on without masks, pose or warping. arXiv:2608.05745, August 2026. arxiv.org/abs/2608.05745
  33. [86]Vidu S2-Editing: real-time video editing adapted to causal streaming. arXiv:2609.11638, September 2026. arxiv.org/abs/2609.11638
  34. [87]Yin et al. From Slow Bidirectional to Fast Autoregressive Video Diffusion Models (CausVid). arXiv:2412.07772. arxiv.org/abs/2412.07772
  35. [88]Decart. MirageLSD: live stream diffusion technical report, July 2025 (archived copy). web.archive.org/web/20250918033439/https://about.decart.ai/publications/mirage
  36. [89]Huang et al. Self Forcing: Bridging the Train-Test Gap in Autoregressive Video Diffusion. arXiv:2506.08009. arxiv.org/abs/2506.08009
  37. [90]Decart. Lucy 2.5: raising the bar for live AI, July 2026. decart.ai/publications/lucy-2-5-raising-the-bar-for-live-ai
  38. [91]ByteDance Seed. Seaweed APT2: autoregressive adversarial post-training for real-time video. seaweed-apt.com/2
  39. [92]NVIDIA. LongLive: Real-time Interactive Long Video Generation. arXiv:2509.22622. arxiv.org/abs/2509.22622
  40. [93]Krea. Krea Realtime 14B. www.krea.ai/blog/krea-realtime-14b
  41. [94]StreamDiffusionV2. MLSys 2026. arXiv:2511.07399. arxiv.org/abs/2511.07399
  42. [95]JD. JoyAI-Video-Edit: real-time instruction video editing (Apache-2.0). arXiv:2608.03974. arxiv.org/abs/2608.03974
  43. [96]Pika. PikaStream 1.0: real-time video chat, April 2026. pika.art/blog/introducing-real-time-video-chat
  44. [97]Wan Team. Wan2.1 open video foundation models (Apache-2.0). github.com/Wan-Video/Wan2.1
  45. [98]Decart. Network requirements for realtime sessions. docs.platform.decart.ai/integrations/network-requirements
  46. [99]Lightricks. LTX-2.x Community License. github.com/Lightricks/LTX-2/blob/main/LICENSE-2_x
  47. [100]Tencent. HunyuanVideo 1.5 licence (Tencent Hunyuan Community License). github.com/Tencent-Hunyuan/HunyuanVideo-1.5/blob/master/LICENSE
  48. [101]Wan Team. Wan2.2 and Wan2.2-Animate (Apache-2.0). huggingface.co/Wan-AI/Wan2.2-Animate-14B
  49. [102]TripVVT: synthetic triplets for video try-on, and the TripVVT-10K dataset (CC BY-NC 4.0). arXiv:2604.27958. arxiv.org/abs/2604.27958
  50. [103]MV-Fashion: multi-view studio fashion video dataset (CC BY-NC-SA 4.0). CVPR 2026. arXiv:2603.08147. arxiv.org/abs/2603.08147