Skip to content
Clothsy AI Talk to us

Chapter 3BThe programme9 min readVersion 1.2 · 30 September 2026

Video and live models

From an offline video try-on model to a live student that draws every frame in 33 to 66 milliseconds on one GPU.

3.8WP6: Video try-on model

Base model and design

The offline video model fine-tunes Wan2.1 14B, which is licensed Apache-2.0 and is the base behind most 2026 video try-on papers [97, 80, 71]. It also serves as the teacher for the live student in WP7. The design follows the pattern the field has converged on:

  • Garment memory. The garment views from WP3 are encoded once and prepended as reference tokens, with a garment-aware position encoding, so the same garment is looked up in every frame [80].
  • No per-frame masks. The model reads the person video directly instead of body masks and pose maps, because those cost more per frame than a distilled generator [85, 84].
  • Keyframes from the image model. For training data and for hard garments, FabricVTON's WP3 image model dresses a few keyframes, and the video model generates the motion around them [81, 84].

Video data engine

No public video dataset can be used commercially, so the data is built from scratch. Real, consented footage is always the target; synthetic re-dressing only ever creates the input [102, 85].

Table 21 Video training data sources and licence position
SourceWhat it providesLicence positionRole
Consented video shootsEach look filmed on three synchronised cameras, a phone and a webcam, with the garment's product shotsOwned, with releases covering AI training and synthetic derivativesReal targets with real garment photos
Licensed motion footagePeople moving in varied settings, without paired garment photosDataset licences [104] or brokered creator footage; free stock sites generally forbid machine-learning use without permission [105]Motion variety for the unpaired stage
Synthetic re-dressed inputsKeyframes re-dressed by an Apache-2.0 image teacher, propagated with Wan2.2-AnimateAll teachers are Apache-2.0 [101]Paired training at scale
Public research setsViViD, VVT, TripVVT-10KNon-commercial or scrapedEvaluation only, never training

Each candidate clip must pass a temporal warping check with RAFT optical flow, frame-to-frame and garment similarity checks, identity and background checks, a vision-language judge, and a human spot check [106]. The annotation tools are chosen for their licences: SAM 2 for segmentation, RTMPose for pose, RAFT for flow and Qwen2.5-VL-7B for captions [107, 108]. OpenPose, Sapiens, CoTracker, SMPL-X and InsightFace are excluded, because their licences forbid commercial use.

Table 22 Video shoot plan
TierPeople and looksGarmentsShoot daysClips produced
Pilot6 adults602Protocol and pipeline tested end to end
Tier A, programme target40 adults, 20 looks each300About 18About 7,000 real clips and 20,000 synthetic pairs, plus 150 hours of licensed footage
Tier B, if funded100 adults, 25 looks each800About 55About 22,000 real clips and 60,000 synthetic pairs, plus 500 hours of licensed footage

Each look follows a fixed motion script of about 100 seconds: standing, a slow turn, walking, arm raises, garment adjustments, sitting and a fast spin. Only consenting adults are filmed, with ID checks. Scale is anchored on published sets of about 9,000 to 10,000 videos.

Training recipe

Table 23 Initial video training configuration
SettingValueBasis
BaseWan2.1 14B, Apache-2.0Used by MagicTryOn, KeyTailor and LiveVVT's teacher
AdapterLoRA first, full fine-tune of attention layers if neededKeyTailor, Eevee
InputsPerson video, garment tokens from several views, optional keyframesMagicTryOn, DreamVVT, BooM-VVT
Clips81 frames, 480p first, then 720pCommon 2026 setting
StagesSynthetic pairs, then real consented pairs, then a high-resolution detail stageMirrors the image recipe
Hardware8 H100 GPUs for about two weeks per full runMagicTryOn and DreamVVT used 8 H20 GPUs

3.9WP7: Live try-on

From offline model to live model

The live model is a 1.3-billion-parameter Wan2.1 student, in its VACE or Fun-Control variant, that learns from the 14B offline model, following the path LiveVVT and Vidu S2 describe [71, 86]. Before any training, the team tests JD's Apache-2.0 JoyAI-Video-Edit zero-shot on garment changes and benchmarks the leading commercial engine as the quality bar [95]. The student is then built in four steps:

  • Causal conversion. The student learns to generate frame by frame from the teacher's outputs, so it never needs future frames [87].
  • Self-forcing and step distillation. It is trained on its own rolled-out frames with a key-value cache, using the Apache-2.0 Self-Forcing code, then distilled to 1 to 4 steps, so errors do not accumulate and each frame is fast [89].
  • Persistent garment memory. The garment photo and description are encoded once per session and kept alongside a frontal keyframe, which keeps a print the same print as the shopper turns [71].
  • Speed engineering. 8-bit arithmetic, compiled kernels and an upscaler from 512p to 720p come after quality is proven.

Streaming system

Figure 2 shows one frame of the loop. The browser sends the camera over WebRTC to a GPU in the nearest region; the model draws the new frame; it streams back to the screen; and the next frame follows about 33 milliseconds later. FabricVTON's existing live page already handles the camera, the consent screen, the session limit and one-time session tokens. Only two small parts of it talk to the rented engine, so FabricVTON's own model can replace it behind the same page. The streaming servers use the open-source LiveKit WebRTC server, and each stream needs about 1.3 to 2.4 megabits per second in each direction [109, 98]. People notice a delay in a mirror-like view at about 150 milliseconds, which open models cannot reach yet, so the design target is about 0.4 seconds from camera to screen. Live serving uses GPUs with hardware video encoders, such as the L40S or RTX PRO 6000, because the H100 and B200 have none.

One frame, repeated about 30 times a secondShopper's deviceBrowser cameraConsent screenSession time limitNo photo storedUpstreamWebRTC videoOne-time session tokenNearest GPU regionGPU: live try-on modelCurrent camera frameGarment memory, set onceIts own recent framesDraws the next frameDownstreamSame person and poseWearing the garmentStreamed backAI-generated labelframe shown on screen; the next one is drawn about 33 ms laterGuardrails in the loopAge and consent check at session startSampled output moderation, watermark and label
Figure 2One frame of live try-on, from camera to screen. The loop repeats about 30 times a second.
Table 24 Live try-on targets against the rented engine
MeasureRented engine todayFabricVTON targetBasis
Resolution720p512p first, 720p through an upscalerLiveVVT 512×384; rented engine 720p
Frame rateUp to 30 fps15 to 20 fps first, then 24Published one-GPU systems reach 11 to 24 fps
First frameAbout 1.5 s, vendor claimUnder 2 sLiveVVT 1.56 s
Cost per minuteUSD 1.20About USD 0.07One H100 per stream at list price
Garment fidelityNot publishedMeasured on the video hold-out, including prints over 60 sWP1 video hold-out
ControlClosed model and termsOwn model, own data, own servingProgramme goal

3.10Evaluation protocol

Table 25 Evaluation measures
MeasureWhat it capturesUse
FID and KIDDistribution-level realismComparability with prior work
SSIM and LPIPSPixel and perceptual similarity in the paired settingComparability and regression tracking
Garment similarity on cropsStructural garment fidelity using self-supervised image featuresAutomatic fidelity proxy
Text and motif accuracyCharacter error rate of printed text and logos, and motif matchingFine-detail fidelity
Garment attribute accuracyHuman labels for fit, length, layer order, closures, print and drapePrimary outcome
Skin-tone driftColour difference (CIEDE2000) on exposed skin before and after try-onFairness
Body-shape distortionChange in keypoints and silhouetteFairness
Pairwise human preferenceOverall quality as judged by raters from several regionsPrimary outcome
Judge agreementKendall τ between automatic judges and human raters, per garment familyTests hypothesis H3
Video realismVFID on full-length clips, with the corrected toolkitComparability for video
Temporal stabilityOptical-flow warping error and VBench consistency scores [110]Flicker and drift
Garment fidelity over timeGarment similarity and text accuracy on frames sampled across each clipPrimary video outcome
Live performanceTime to first frame, sustained frame rate, 95th-percentile frame latency, drift per 10 secondsLive readiness
Latency and costSeconds per photo and USD per image or per stream-minuteProduct readiness

The primary outcomes, human-rated garment fidelity and pairwise preference, are fixed before experiments begin. Results carry bootstrap confidence intervals and are broken out by garment family, presentation, photo or video condition, skin-tone group and body shape.

3.11Research questions and hypotheses

Table 26 Research questions and hypotheses
IDResearch questionHypothesisTested in
H1How do current systems fail on garments beyond stitched studio wear?Failures are specific and measurable, concentrated in layer order, multi-piece consistency, fine print, back views and drapePaper 1
H2Are the failures uneven across skin tones and body shapes?Skin-tone drift and body distortion are larger for darker skin tones and larger bodiesPaper 1
H3Do standard metrics and automatic judges detect these failures?Metrics and judges calibrated on stitched studio garments agree less with human raters on the harder garment families and on real shoppers' photosPaper 1
H4Can multi-view conditioning and pseudo-pair adaptation close the gap?Multi-view garments plus fine-tuning on pseudo-pairs close most of the gap with a few thousand real pairsPaper 2
H5Can an image try-on model teach a video model?Video pairs made by re-dressing real clips with the image model, keeping the real clip as the target, train a video model that beats open video baselines without scraped dataPaper 3
H6Can garment fidelity survive the move to live generation?A causal student with a persistent garment memory and self-forcing training keeps prints and logos stable for 60 seconds or more at 15 fps or betterPaper 4

3.12Compute estimate

Table 27 Estimated compute budget
ItemBasisEstimate (USD)
Baseline auditAbout 3,000 pairs across ten systems, open models on cloud GPUs plus paid APIs at USD 0.04 to 0.075 per image1,000 to 2,000
Synthetic image triplets500,000 teacher outputs at 6.5 seconds each on an L40S at USD 2.27 per hourAbout 2,000
Image LoRA runs and ablationsLayering VTON scale, 300 to 600 H200 GPU-hours per full run, plus five to eight smaller runs6,500 to 13,000
Image distillation and serving testsStudent training plus speed-laboratory runs1,000 to 3,000
Video annotation and filteringAbout 300 to 800 GPU-hours for segmentation, pose, flow and captions1,000 to 3,000
Synthetic video pairsAbout 2,000 to 3,000 GPU-hours, including over-generation before filtering6,000 to 12,000
Offline video modelThree two-week runs on 8 H100s (about 2,700 GPU-hours each at USD 3.95) plus ablations30,000 to 45,000
Live distillation1,500 to 5,000 H100-hours across several attempts; published recipes used about 100 to 3,0007,000 to 20,000
Streaming tests and live pilotServing GPUs, TURN relays and load tests3,000 to 6,000
Total computePhotography, video shoots, raters and annotation are budgeted separately60,000 to 105,000

These are planning estimates. The L40S rate and 6.5-second time come from FabricVTON's measurements; H100 and H200 rates are cloud list prices; training hours are extrapolated from published recipes. Cloud credits would reduce the cash cost.

Sources in this chapter

  1. [7]FabricVTON. Internal technical documentation: What Is Built (updated 18 and 25 September 2026) and Project Status (24 September 2026). Available to reviewers on request.
  2. [71]LiveVVT: High-Fidelity Video Virtual Try-On in Real Time. arXiv:2608.26714, August 2026. arxiv.org/abs/2608.26714
  3. [80]MagicTryOn: video virtual try-on on Wan2.1. arXiv:2505.21325, 2025. arxiv.org/abs/2505.21325
  4. [81]DreamVVT: keyframe-first video virtual try-on. ByteDance. arXiv:2508.02807, 2025. arxiv.org/abs/2508.02807
  5. [84]BooM-VVT: mask-free video try-on with garment-sensitive keyframes. ACM MM 2026. arXiv:2609.04120. arxiv.org/abs/2609.04120
  6. [85]UniVVT: video try-on without masks, pose or warping. arXiv:2608.05745, August 2026. arxiv.org/abs/2608.05745
  7. [86]Vidu S2-Editing: real-time video editing adapted to causal streaming. arXiv:2609.11638, September 2026. arxiv.org/abs/2609.11638
  8. [87]Yin et al. From Slow Bidirectional to Fast Autoregressive Video Diffusion Models (CausVid). arXiv:2412.07772. arxiv.org/abs/2412.07772
  9. [89]Huang et al. Self Forcing: Bridging the Train-Test Gap in Autoregressive Video Diffusion. arXiv:2506.08009. arxiv.org/abs/2506.08009
  10. [95]JD. JoyAI-Video-Edit: real-time instruction video editing (Apache-2.0). arXiv:2608.03974. arxiv.org/abs/2608.03974
  11. [97]Wan Team. Wan2.1 open video foundation models (Apache-2.0). github.com/Wan-Video/Wan2.1
  12. [98]Decart. Network requirements for realtime sessions. docs.platform.decart.ai/integrations/network-requirements
  13. [101]Wan Team. Wan2.2 and Wan2.2-Animate (Apache-2.0). huggingface.co/Wan-AI/Wan2.2-Animate-14B
  14. [102]TripVVT: synthetic triplets for video try-on, and the TripVVT-10K dataset (CC BY-NC 4.0). arXiv:2604.27958. arxiv.org/abs/2604.27958
  15. [104]Shutterstock. Data licensing and the contributor fund. submit.shutterstock.com/help/en/articles/10594694-shutterstock-data-licensing-and-the-contributor-fund
  16. [105]Pexels. Terms of service, including limits on machine-learning use. www.pexels.com/terms-of-service
  17. [106]Teed and Deng. RAFT optical flow (BSD-3-Clause). github.com/princeton-vl/RAFT
  18. [107]Meta. SAM 2: Segment Anything in Images and Videos (Apache-2.0). github.com/facebookresearch/sam2
  19. [108]OpenMMLab. MMPose and RTMPose (Apache-2.0). github.com/open-mmlab/mmpose
  20. [109]LiveKit. Open-source WebRTC media server. github.com/livekit/livekit
  21. [110]VBench: comprehensive benchmark suite for video generative models. github.com/Vchitect/VBench