This part of the docs summarises the state of image virtual try-on as of September 2026. It draws on published papers, model cards, licence files and company disclosures. Where a figure could not be verified from a primary source, the text says so.
2.1The task
An image try-on model takes a photograph of a person and a photograph of a garment, and produces a photograph of that person wearing that garment. The garment photo may be a product shot or a photo of someone else wearing it. Two evaluation settings are standard. In the paired setting, the model re-dresses a person in the garment they already wear, so the output can be compared with the real photograph. In the unpaired setting the garment is new, so only distribution-level realism can be scored.
Models also differ in what else they need. Older methods require a mask of the region to repaint, a pose map and a body-part segmentation. Newer mask-free methods take only the two photographs, which makes them easier to deploy and less brittle when segmentation fails.
Video and live try-on. Video try-on takes a video of a person and a garment image, and returns the video with the person wearing the garment. It must keep identity, motion and background, and keep the garment identical from frame to frame. Live try-on does the same on a camera stream as it arrives. Each frame must be ready in about 33 to 66 milliseconds, drawn only from the current camera frame and what the model has already produced, because future frames do not exist yet.
2.2How the architectures evolved
- 2023Warp the garment, then paintTryOnDiffusion, LaDI-VTON
- 2024Garment reference network sharing attentionStableVITON, OOTDiffusion, IDM-VTON, Leffa
- Late 2024One network, images side by sideCatVTON, TPD
- 2025Diffusion transformer with garment tokens in contextFitDiT, OmniTry, Voost, FASHN VTON 1.5
- 2026LoRA fine-tuning of general image editorsLayering VTON, TAMF-VTON, RealFit
- 2026Fashion foundation models with reinforcement learningTstars-Tryon, Oxygen-TryOn
| Period | Approach | Representative work | Main limitation |
|---|---|---|---|
| 2023 | Warp the garment, then paint | TryOnDiffusion [31], LaDI-VTON [32] | Explicit warping fails on large deformation. TryOnDiffusion needed about 4 million private pairs. |
| 2024 | Garment reference network sharing attention | StableVITON [33], OOTDiffusion, IDM-VTON, Leffa | Needs masks, pose and parsing. Built on older latent backbones that blur fine detail. |
| Late 2024 | One network, images side by side | CatVTON, TPD | Lightweight, but limited by the older backbone. |
| 2025 | Diffusion transformer with garment tokens in context | FitDiT, OmniTry [34], Voost [35], FASHN VTON 1.5 | Mask-free training relies on synthetic pairs made by an earlier model. |
| 2026 | LoRA fine-tuning of general image-editing models | Layering VTON, TAMF-VTON [36] on Qwen-Image-Edit; RealFit [37] on FLUX Kontext | Base-model licences vary. Large models are slow and costly to serve. |
| 2026 | Fashion foundation models with reinforcement learning | Tstars-Tryon (5B), Oxygen-TryOn (16B) | Require tens of millions of images and in-house reward models. |
OOTDiffusion, often cited as the reference open model, belongs to the 2024 generation. It pairs a Stable Diffusion 1.5 denoiser with an outfitting network whose features are fused through self-attention, and it reports results at 512×384 [25]. Its fusion idea survives in newer designs. Its backbone, resolution and CC BY-NC-SA 4.0 licence make it unsuitable as the foundation for a commercial model today.
2.3Reported results on the standard benchmark
Table 6 lists results on VITON-HD as reported in each paper. The numbers are indicative only. Papers differ in resolution, FID implementation and whether baselines were re-run, so differences below about one FID point are within that noise.
| Method | Year | Base | FID (unpaired) ↓ | SSIM ↑ | LPIPS ↓ | Weights licence |
|---|---|---|---|---|---|---|
| LaDI-VTON | 2023 | SD 2 | 9.41 | 0.876 | 0.091 | CC BY-NC |
| OOTDiffusion | 2024 | SD 1.5 | 8.81 * | 0.878 | 0.071 | CC BY-NC-SA |
| IDM-VTON | 2024 | SDXL | 6.29 † | 0.870 | 0.102 | CC BY-NC-SA |
| CatVTON | 2025 | SD 1.5 | 9.02 | 0.870 | 0.057 | CC BY-NC-SA |
| Leffa | 2025 | SD 1.5 | 8.52 | 0.899 | 0.048 | MIT code, non-commercial data |
| FitDiT | 2024 | SD 3 | 8.20 | 0.899 | 0.066 | CC BY-NC-SA |
| Voost | 2025 | DiT | 8.98 | 0.898 | 0.056 | Not released |
| FastFit [38] | 2025 | SD 1.5 | 8.63 | 0.885 | 0.078 | Non-commercial |
| CORAL [39] | 2026 | FLUX Fill | 8.76 | 0.907 | 0.048 | Non-commercial base |
| RealFit | 2026 | FLUX Kontext | 7.74 | 0.907 | 0.049 | Non-commercial base |
| TAMF-VTON | 2026 | Qwen-Image-Edit | 6.27 | 0.913 | 0.052 | No code released |
| Oxygen-TryOn | 2026 | JoyAI-Image-Edit, 16B | 9.15 ‡ | 0.914 | 0.050 | Weights announced |
| FASHN VTON 1.5 [40] | 2026 | Own 972M, pixel space | Not published | Not published | Not published | Apache-2.0 |
* Reported at 512×384. † Evaluation setting not stated; CatVTON's re-run of IDM-VTON gives 9.84. ‡ Oxygen-TryOn re-evaluated all baselines itself and reports the best paired FID, 3.94. SSIM and LPIPS are paired-setting values.
Two conclusions follow. First, VITON-HD is saturating: the spread among modern methods is small relative to measurement noise, and the benchmark itself covers only upper-body garments. Second, general image editors are now serious competitors. On VTEdit-Bench, published at ECCV 2026, FLUX.2 models ranked best overall among general editors and Qwen-Image-Edit-2511 came close [41]. In Alibaba's human study, Tstars-Tryon beat Google's Nano Banana Pro in 41.1% of comparisons, tied in 41.6% and lost 17.3% [3].
2.4What distinguishes the strongest systems
- Scale and synthetic pairs. Google trained TryOnDiffusion on about 4 million pairs. FASHN trained VTON 1.5 on 18 million masked pairs and then 4 million synthetic triplets, and JD built Oxygen-TryOn from more than 50 million raw images [31, 40, 15]. A synthetic triplet keeps a real photograph as the target and uses an earlier model to create the input. The model therefore learns to produce real photographs without needing a mask.
- Editing foundation models plus LoRA. Layering VTON fine-tuned Qwen-Image-Edit with LoRA of rank 32 on the attention projections, on one NVIDIA H200 GPU: 20,000 steps on 29,151 pairs, then 5,000 steps on 5,768 pairs [16]. TAMF-VTON reports the lowest unpaired FID in Table 6 with a mixture-of-experts LoRA on the same base [36].
- Preference optimisation. Tstars-Tryon and Oxygen-TryOn follow pre-training and supervised fine-tuning with reinforcement learning. Oxygen's reward combines an 8-billion-parameter vision-language model trained on 100,000 preference pairs with a Gemini rubric judge [15]. TryOnReward released about 100,000 human ratings for this purpose [42].
- Distillation and caching for speed. FastFit caches garment features for a 3.5× speed-up, DirectTryOn generates in one step in 0.48 seconds, and Tstars-Tryon serves in about 3.9 seconds after step distillation [38, 43, 3].
- Small specialists remain competitive. FASHN VTON 1.5 generates directly in pixel space with 972 million parameters and no autoencoder. FASHN states it can be trained from scratch for USD 5,000 to 10,000 [30, 44]. Pixel-space generation avoids the eight-fold compression that blurs fine print.
2.5Industrial landscape
| Organisation | System | Public facts on architecture and data | Deployment |
|---|---|---|---|
| TryOnDiffusion; Search try-on; Vertex virtual-try-on-001 | Two-UNet cascade to 1024 pixels, about 4 million pairs; later described as a custom fashion image model | US Search from July 2025; UK and India from December 2025; enterprise API with SynthID watermarking [14] | |
| FASHN AI | VTON 1.5 open weights; API v1.6; Try-On Max | 972M pixel-space transformer; 18 million pairs plus 4 million synthetic triplets | API at 864×1296; USD 0.075 per image on demand [45] |
| Alibaba (Taobao) | Tstars-Tryon 1.0 | 5B transformer; pre-training, fine-tuning and reinforcement learning; up to six references | Taobao app, several million users, about 3.9 seconds |
| JD.com | Oxygen-TryOn | 16B transformer; 50 million raw images; learned reward plus a vision-language judge | Weights announced as forthcoming |
| ByteDance | Seedream 4.0 | General editor marketed for multi-garment try-on at up to 4K | API [46] |
| Inditex | Zara Try-On | Avatar from the shopper's photos; vendor not disclosed | More than 7 million sessions in 43 markets |
| Meesho | Saree try-on with Google Cloud | Warping, 3D rendering and Imagen, announced in 2024 | Announced as coming soon [47] |
Industrial try-on is now deployed across North America, Europe and Asia. This review found no public disclosure from these organisations of try-on quality broken down by skin tone or body shape, and no deployed generative try-on built for draped garments: Meesho's announced saree try-on relied on warping and 3D rendering rather than a generative try-on model [47].
2.6Open foundation models and their licences
For a company, the licence of the base model decides what can be sold. A LoRA inherits its base model's terms, and some licences also restrict using a model's outputs to train other models.
| Model | Released | Size | Person and garment as separate inputs | Licence | Commercial use |
|---|---|---|---|---|---|
| Qwen-Image-Edit-2511 [48] | Dec 2025 | 20B MMDiT | Yes, native multi-image | Apache-2.0 | Yes |
| FLUX.2 klein base 4B [49] | Jan 2026 | 4B | Yes, multi-reference | Apache-2.0 | Yes |
| FASHN VTON 1.5 | Jan 2026 | 972M | Built for try-on | Apache-2.0; parser non-commercial | Yes, after parser replacement |
| LongCat-Image-Edit [50] | Dec 2025 | 6B | One image; use a side-by-side canvas | Apache-2.0 | Yes |
| HiDream-O1-Image [51] | May 2026 | 8B, pixel space | One reference for editing | MIT | Yes |
| Qwen-Image-2.1 | Sep 2026 | 7B | Up to 10 references | Qwen Research License | No |
| FLUX.1 Kontext dev, FLUX.2 dev, klein 9B | 2025 to 2026 | 12B, 32B, 9B | Yes | FLUX Non-Commercial | No |
| HunyuanImage 3.0 | Jan 2026 | 80B MoE | Up to 3 images | Tencent community licence | Restricted by region; outputs may not improve other models |
| Stable Diffusion 3.5 | Oct 2024 | 8.1B and 2.5B | No native editing | Stability Community | Only under USD 1M revenue |
2.7Datasets and their licences
| Dataset | Content | Commercial training |
|---|---|---|
| VITON-HD | 13,679 upper-body pairs at 1024×768 | No. CC BY-NC 4.0 |
| DressCode | About 54,000 garments across three categories | No. Not released to private companies |
| DeepFashion, DeepFashion-MultiModal | About 800,000 and 44,000 images | No. Research use only |
| StreetTryOn | About 12,400 street images | No. Commercial use prohibited |
| SHHQ | About 40,000 full-body images | No. Research use only |
| IGPair | More than 300,000 pairs | No. Academic and personal use |
| IndoFashion | 106,000 images in 15 Indian ethnic categories, for classification | Unclear. Images gathered from e-commerce and web search |
| BD-VITON | 1,013 pairs of Bangladeshi garments | No clear commercial terms. The paper is CC BY-NC-SA 4.0; no separate dataset licence is stated |
| Garments2Look [52] | About 80,000 multi-garment outfits | Risky. About half synthesised with a commercial API, the rest from web images |
None of these datasets clearly permits commercial training. Evaluation on the public datasets is possible only where their terms allow research use. A commercial model must be trained on data that FabricVTON owns or licenses.
2.8How try-on is evaluated, and where evaluation falls short
FID and KID measure whether outputs look like real photographs, not whether the right garment was transferred. SSIM and LPIPS need a ground-truth photograph and are dominated by the background. OpenVTON-Bench measured agreement with human rankings at Kendall τ of 0.611 for SSIM, against 0.833 for its own metric combining a vision-language judge and segmentation [53].
Several try-on benchmarks appeared in 2025 and 2026: VTBench, VTONQA, OpenVTON-Bench, VTON-QBench with 431,800 human annotations, VTEdit-Bench and TryOnReward [54, 55, 56]. They improve on classic metrics, but none covers draped garments, and none reports results by skin tone or body shape. A new general-purpose benchmark would not be novel on its own. One that spans garments from many regions, includes real shoppers' photos, is stratified by skin tone and body shape, and is built from consented, licensed photographs would be.
Sources in this chapter
- [3]Alibaba Taobao. Tstars-Tryon 1.0. arXiv:2604.19748, April 2026. arxiv.org/abs/2604.19748
- [4]Choi et al. VITON-HD: High-Resolution Virtual Try-On. CVPR 2021. Dataset licence CC BY-NC 4.0. github.com/shadow2496/VITON-HD
- [5]Morelli et al. Dress Code: High-Resolution Multi-Category Virtual Try-On. ECCV 2022. Dataset access terms. github.com/aimagelab/dress-code
- [6]Virtual Try-On for Cultural Clothing: A Benchmarking Study (BD-VITON). arXiv:2603.07291, March 2026. Paper licensed CC BY-NC-SA 4.0; no separate dataset licence stated. arxiv.org/abs/2603.07291
- [14]Google Cloud. Virtual Try-On on Vertex AI (virtual-try-on-001). Documentation. docs.cloud.google.com/vertex-ai/generative-ai/docs/models/imagen/virtual-try-on-preview-08-04
- [15]JD.com. Oxygen-TryOn. arXiv:2607.21694, July 2026. arxiv.org/abs/2607.21694
- [16]Feng, Chen, Shan and Kemelmacher-Shlizerman. Layering Virtual Try-On. ECCV 2026. arXiv:2607.22924. arxiv.org/abs/2607.22924
- [17]Zhou et al. Learning Flow Fields in Attention for Controllable Person Image Generation (Leffa). CVPR 2025. arXiv:2412.08486. arxiv.org/abs/2412.08486
- [18]Chong et al. CatVTON: Concatenation Is All You Need for Virtual Try-On with Diffusion Models. ICLR 2025. arXiv:2407.15886. arxiv.org/abs/2407.15886
- [19]Rajput and Aneja. IndoFashion: Apparel Classification for Indian Ethnic Clothes. CVPR Workshops 2021. arXiv:2104.02830. arxiv.org/abs/2104.02830
- [25]Xu et al. OOTDiffusion: Outfitting Fusion Based Latent Diffusion for Controllable Virtual Try-On. AAAI 2025. arXiv:2403.01779. arxiv.org/abs/2403.01779
- [26]Choi et al. Improving Diffusion Models for Authentic Virtual Try-On in the Wild (IDM-VTON). ECCV 2024. arXiv:2403.05139. arxiv.org/abs/2403.05139
- [27]Jiang et al. FitDiT: Advancing the Authentic Garment Details for High-Fidelity Virtual Try-On. arXiv:2411.10499. arxiv.org/abs/2411.10499
- [28]Qwen Team. Qwen-Image-2.1 licence (Qwen Research License Agreement), September 2026. github.com/QwenLM/Qwen-Image-2.1/blob/main/LICENSE
- [29]Black Forest Labs. FLUX.2 [klein] 9B licence (FLUX Non-Commercial License). huggingface.co/black-forest-labs/FLUX.2-klein-9B/blob/main/LICENSE.md
- [30]FASHN AI. FASHN VTON v1.5: Efficient Maskless Virtual Try-On in Pixel Space. Research page. fashn.ai/research/vton-1-5
- [31]Zhu et al. TryOnDiffusion: A Tale of Two UNets. CVPR 2023. arXiv:2306.08276. arxiv.org/abs/2306.08276
- [32]Morelli et al. LaDI-VTON: Latent Diffusion Textual-Inversion Enhanced Virtual Try-On. ACM MM 2023. arXiv:2305.13501. arxiv.org/abs/2305.13501
- [33]Kim et al. StableVITON: Learning Semantic Correspondence with Latent Diffusion Model for Virtual Try-On. CVPR 2024. arXiv:2312.01725. arxiv.org/abs/2312.01725
- [34]OmniTry: mask-free virtual try-on for wearable objects. arXiv:2508.13632. arxiv.org/abs/2508.13632
- [35]Voost: a diffusion transformer for joint virtual try-on and try-off. SIGGRAPH Asia 2025. arXiv:2508.04825. arxiv.org/abs/2508.04825
- [36]TAMF-VTON. arXiv:2607.14807, July 2026. arxiv.org/abs/2607.14807
- [37]RealFit. arXiv:2609.25881, September 2026. arxiv.org/abs/2609.25881
- [38]FastFit: cacheable multi-reference virtual try-on. arXiv:2508.20586. arxiv.org/abs/2508.20586
- [39]CORAL: correspondence-aligned virtual try-on. arXiv:2602.17636, 2026. arxiv.org/abs/2602.17636
- [40]FASHN AI. fashn-vton-1.5 model card, Hugging Face. huggingface.co/fashn-ai/fashn-vton-1.5
- [41]VTEdit-Bench. ECCV 2026. arXiv:2603.11734. arxiv.org/abs/2603.11734
- [42]TryOnReward. arXiv:2609.13259, September 2026. arxiv.org/abs/2609.13259
- [43]DirectTryOn. arXiv:2605.12939, May 2026. arxiv.org/abs/2605.12939
- [44]FASHN AI. Open-sourcing FASHN VTON v1.5. Blog post, 2026. fashn.ai/blog/fashn-vton-1-5-open-source-release
- [45]FASHN AI. API pricing. Help centre. help.fashn.ai/plans-and-pricing/api-pricing
- [46]ByteDance Seed. Seedream 4.0 officially released. September 2025. seed.bytedance.com/en/blog/seedream-4-0-officially-released-beyond-drawing-into-imagination
- [47]Google Cloud. Virtual try-on technology with Google Cloud AI (Meesho case). July 2024. cloud.google.com/blog/products/ai-machine-learning/virtual-try-on-technology-with-google-cloud-ai
- [48]Qwen Team. Qwen-Image-Edit-2511 model card (Apache-2.0). huggingface.co/Qwen/Qwen-Image-Edit-2511
- [49]Black Forest Labs. FLUX.2 [klein] base 4B model card (Apache-2.0). huggingface.co/black-forest-labs/FLUX.2-klein-base-4B
- [50]Meituan LongCat. LongCat-Image-Edit model card. huggingface.co/meituan-longcat/LongCat-Image-Edit
- [51]HiDream. HiDream-O1-Image model card (MIT). huggingface.co/HiDream-ai/HiDream-O1-Image
- [52]Garments2Look. CVPR 2026. arXiv:2603.14153. arxiv.org/abs/2603.14153
- [53]OpenVTON-Bench. arXiv:2601.22725, January 2026. arxiv.org/abs/2601.22725
- [54]VTBench: a hierarchical virtual try-on benchmark. arXiv:2505.19571, May 2025. arxiv.org/abs/2505.19571
- [55]VTONQA. arXiv:2601.02945, January 2026. arxiv.org/abs/2601.02945
- [56]VTON-IQA and VTON-QBench. arXiv:2603.13057, March 2026. arxiv.org/abs/2603.13057
