Skip to content
Clothsy AI Talk to us

Chapter 2AState of the art11 min readVersion 1.2 · 30 September 2026

Image try-on today

How image try-on is built in 2026, which models and data a company can legally use, and where evaluation falls short.

This part of the docs summarises the state of image virtual try-on as of September 2026. It draws on published papers, model cards, licence files and company disclosures. Where a figure could not be verified from a primary source, the text says so.

2.1The task

An image try-on model takes a photograph of a person and a photograph of a garment, and produces a photograph of that person wearing that garment. The garment photo may be a product shot or a photo of someone else wearing it. Two evaluation settings are standard. In the paired setting, the model re-dresses a person in the garment they already wear, so the output can be compared with the real photograph. In the unpaired setting the garment is new, so only distribution-level realism can be scored.

Models also differ in what else they need. Older methods require a mask of the region to repaint, a pose map and a body-part segmentation. Newer mask-free methods take only the two photographs, which makes them easier to deploy and less brittle when segmentation fails.

Video and live try-on. Video try-on takes a video of a person and a garment image, and returns the video with the person wearing the garment. It must keep identity, motion and background, and keep the garment identical from frame to frame. Live try-on does the same on a camera stream as it arrives. Each frame must be ready in about 33 to 66 milliseconds, drawn only from the current camera frame and what the model has already produced, because future frames do not exist yet.

2.2How the architectures evolved

  1. 2023Warp the garment, then paintTryOnDiffusion, LaDI-VTON
  2. 2024Garment reference network sharing attentionStableVITON, OOTDiffusion, IDM-VTON, Leffa
  3. Late 2024One network, images side by sideCatVTON, TPD
  4. 2025Diffusion transformer with garment tokens in contextFitDiT, OmniTry, Voost, FASHN VTON 1.5
  5. 2026LoRA fine-tuning of general image editorsLayering VTON, TAMF-VTON, RealFit
  6. 2026Fashion foundation models with reinforcement learningTstars-Tryon, Oxygen-TryOn
Figure 2.1Six generations of virtual try-on architecture in four years. Each generation kept the best idea of the last and moved to a stronger backbone.
Table 5 Generations of virtual try-on architecture
PeriodApproachRepresentative workMain limitation
2023Warp the garment, then paintTryOnDiffusion [31], LaDI-VTON [32]Explicit warping fails on large deformation. TryOnDiffusion needed about 4 million private pairs.
2024Garment reference network sharing attentionStableVITON [33], OOTDiffusion, IDM-VTON, LeffaNeeds masks, pose and parsing. Built on older latent backbones that blur fine detail.
Late 2024One network, images side by sideCatVTON, TPDLightweight, but limited by the older backbone.
2025Diffusion transformer with garment tokens in contextFitDiT, OmniTry [34], Voost [35], FASHN VTON 1.5Mask-free training relies on synthetic pairs made by an earlier model.
2026LoRA fine-tuning of general image-editing modelsLayering VTON, TAMF-VTON [36] on Qwen-Image-Edit; RealFit [37] on FLUX KontextBase-model licences vary. Large models are slow and costly to serve.
2026Fashion foundation models with reinforcement learningTstars-Tryon (5B), Oxygen-TryOn (16B)Require tens of millions of images and in-house reward models.

OOTDiffusion, often cited as the reference open model, belongs to the 2024 generation. It pairs a Stable Diffusion 1.5 denoiser with an outfitting network whose features are fused through self-attention, and it reports results at 512×384 [25]. Its fusion idea survives in newer designs. Its backbone, resolution and CC BY-NC-SA 4.0 licence make it unsuitable as the foundation for a commercial model today.

2.3Reported results on the standard benchmark

Table 6 lists results on VITON-HD as reported in each paper. The numbers are indicative only. Papers differ in resolution, FID implementation and whether baselines were re-run, so differences below about one FID point are within that noise.

Table 6 Reported VITON-HD results for selected methods
MethodYearBaseFID (unpaired) ↓SSIM ↑LPIPS ↓Weights licence
LaDI-VTON2023SD 29.410.8760.091CC BY-NC
OOTDiffusion2024SD 1.58.81 *0.8780.071CC BY-NC-SA
IDM-VTON2024SDXL6.29 †0.8700.102CC BY-NC-SA
CatVTON2025SD 1.59.020.8700.057CC BY-NC-SA
Leffa2025SD 1.58.520.8990.048MIT code, non-commercial data
FitDiT2024SD 38.200.8990.066CC BY-NC-SA
Voost2025DiT8.980.8980.056Not released
FastFit [38]2025SD 1.58.630.8850.078Non-commercial
CORAL [39]2026FLUX Fill8.760.9070.048Non-commercial base
RealFit2026FLUX Kontext7.740.9070.049Non-commercial base
TAMF-VTON2026Qwen-Image-Edit6.270.9130.052No code released
Oxygen-TryOn2026JoyAI-Image-Edit, 16B9.15 ‡0.9140.050Weights announced
FASHN VTON 1.5 [40]2026Own 972M, pixel spaceNot publishedNot publishedNot publishedApache-2.0

* Reported at 512×384. † Evaluation setting not stated; CatVTON's re-run of IDM-VTON gives 9.84. ‡ Oxygen-TryOn re-evaluated all baselines itself and reports the best paired FID, 3.94. SSIM and LPIPS are paired-setting values.

Two conclusions follow. First, VITON-HD is saturating: the spread among modern methods is small relative to measurement noise, and the benchmark itself covers only upper-body garments. Second, general image editors are now serious competitors. On VTEdit-Bench, published at ECCV 2026, FLUX.2 models ranked best overall among general editors and Qwen-Image-Edit-2511 came close [41]. In Alibaba's human study, Tstars-Tryon beat Google's Nano Banana Pro in 41.1% of comparisons, tied in 41.6% and lost 17.3% [3].

2.4What distinguishes the strongest systems

  • Scale and synthetic pairs. Google trained TryOnDiffusion on about 4 million pairs. FASHN trained VTON 1.5 on 18 million masked pairs and then 4 million synthetic triplets, and JD built Oxygen-TryOn from more than 50 million raw images [31, 40, 15]. A synthetic triplet keeps a real photograph as the target and uses an earlier model to create the input. The model therefore learns to produce real photographs without needing a mask.
  • Editing foundation models plus LoRA. Layering VTON fine-tuned Qwen-Image-Edit with LoRA of rank 32 on the attention projections, on one NVIDIA H200 GPU: 20,000 steps on 29,151 pairs, then 5,000 steps on 5,768 pairs [16]. TAMF-VTON reports the lowest unpaired FID in Table 6 with a mixture-of-experts LoRA on the same base [36].
  • Preference optimisation. Tstars-Tryon and Oxygen-TryOn follow pre-training and supervised fine-tuning with reinforcement learning. Oxygen's reward combines an 8-billion-parameter vision-language model trained on 100,000 preference pairs with a Gemini rubric judge [15]. TryOnReward released about 100,000 human ratings for this purpose [42].
  • Distillation and caching for speed. FastFit caches garment features for a 3.5× speed-up, DirectTryOn generates in one step in 0.48 seconds, and Tstars-Tryon serves in about 3.9 seconds after step distillation [38, 43, 3].
  • Small specialists remain competitive. FASHN VTON 1.5 generates directly in pixel space with 972 million parameters and no autoencoder. FASHN states it can be trained from scratch for USD 5,000 to 10,000 [30, 44]. Pixel-space generation avoids the eight-fold compression that blurs fine print.

2.5Industrial landscape

Table 7 How leading companies build and deploy try-on
OrganisationSystemPublic facts on architecture and dataDeployment
GoogleTryOnDiffusion; Search try-on; Vertex virtual-try-on-001Two-UNet cascade to 1024 pixels, about 4 million pairs; later described as a custom fashion image modelUS Search from July 2025; UK and India from December 2025; enterprise API with SynthID watermarking [14]
FASHN AIVTON 1.5 open weights; API v1.6; Try-On Max972M pixel-space transformer; 18 million pairs plus 4 million synthetic tripletsAPI at 864×1296; USD 0.075 per image on demand [45]
Alibaba (Taobao)Tstars-Tryon 1.05B transformer; pre-training, fine-tuning and reinforcement learning; up to six referencesTaobao app, several million users, about 3.9 seconds
JD.comOxygen-TryOn16B transformer; 50 million raw images; learned reward plus a vision-language judgeWeights announced as forthcoming
ByteDanceSeedream 4.0General editor marketed for multi-garment try-on at up to 4KAPI [46]
InditexZara Try-OnAvatar from the shopper's photos; vendor not disclosedMore than 7 million sessions in 43 markets
MeeshoSaree try-on with Google CloudWarping, 3D rendering and Imagen, announced in 2024Announced as coming soon [47]

Industrial try-on is now deployed across North America, Europe and Asia. This review found no public disclosure from these organisations of try-on quality broken down by skin tone or body shape, and no deployed generative try-on built for draped garments: Meesho's announced saree try-on relied on warping and 3D rendering rather than a generative try-on model [47].

2.6Open foundation models and their licences

For a company, the licence of the base model decides what can be sold. A LoRA inherits its base model's terms, and some licences also restrict using a model's outputs to train other models.

Table 8 Candidate base models for a commercial try-on model
ModelReleasedSizePerson and garment as separate inputsLicenceCommercial use
Qwen-Image-Edit-2511 [48]Dec 202520B MMDiTYes, native multi-imageApache-2.0Yes
FLUX.2 klein base 4B [49]Jan 20264BYes, multi-referenceApache-2.0Yes
FASHN VTON 1.5Jan 2026972MBuilt for try-onApache-2.0; parser non-commercialYes, after parser replacement
LongCat-Image-Edit [50]Dec 20256BOne image; use a side-by-side canvasApache-2.0Yes
HiDream-O1-Image [51]May 20268B, pixel spaceOne reference for editingMITYes
Qwen-Image-2.1Sep 20267BUp to 10 referencesQwen Research LicenseNo
FLUX.1 Kontext dev, FLUX.2 dev, klein 9B2025 to 202612B, 32B, 9BYesFLUX Non-CommercialNo
HunyuanImage 3.0Jan 202680B MoEUp to 3 imagesTencent community licenceRestricted by region; outputs may not improve other models
Stable Diffusion 3.5Oct 20248.1B and 2.5BNo native editingStability CommunityOnly under USD 1M revenue

2.7Datasets and their licences

Table 9 Try-on datasets and whether they permit commercial training
DatasetContentCommercial training
VITON-HD13,679 upper-body pairs at 1024×768No. CC BY-NC 4.0
DressCodeAbout 54,000 garments across three categoriesNo. Not released to private companies
DeepFashion, DeepFashion-MultiModalAbout 800,000 and 44,000 imagesNo. Research use only
StreetTryOnAbout 12,400 street imagesNo. Commercial use prohibited
SHHQAbout 40,000 full-body imagesNo. Research use only
IGPairMore than 300,000 pairsNo. Academic and personal use
IndoFashion106,000 images in 15 Indian ethnic categories, for classificationUnclear. Images gathered from e-commerce and web search
BD-VITON1,013 pairs of Bangladeshi garmentsNo clear commercial terms. The paper is CC BY-NC-SA 4.0; no separate dataset licence is stated
Garments2Look [52]About 80,000 multi-garment outfitsRisky. About half synthesised with a commercial API, the rest from web images

None of these datasets clearly permits commercial training. Evaluation on the public datasets is possible only where their terms allow research use. A commercial model must be trained on data that FabricVTON owns or licenses.

2.8How try-on is evaluated, and where evaluation falls short

FID and KID measure whether outputs look like real photographs, not whether the right garment was transferred. SSIM and LPIPS need a ground-truth photograph and are dominated by the background. OpenVTON-Bench measured agreement with human rankings at Kendall τ of 0.611 for SSIM, against 0.833 for its own metric combining a vision-language judge and segmentation [53].

Several try-on benchmarks appeared in 2025 and 2026: VTBench, VTONQA, OpenVTON-Bench, VTON-QBench with 431,800 human annotations, VTEdit-Bench and TryOnReward [54, 55, 56]. They improve on classic metrics, but none covers draped garments, and none reports results by skin tone or body shape. A new general-purpose benchmark would not be novel on its own. One that spans garments from many regions, includes real shoppers' photos, is stratified by skin tone and body shape, and is built from consented, licensed photographs would be.

Sources in this chapter

  1. [3]Alibaba Taobao. Tstars-Tryon 1.0. arXiv:2604.19748, April 2026. arxiv.org/abs/2604.19748
  2. [4]Choi et al. VITON-HD: High-Resolution Virtual Try-On. CVPR 2021. Dataset licence CC BY-NC 4.0. github.com/shadow2496/VITON-HD
  3. [5]Morelli et al. Dress Code: High-Resolution Multi-Category Virtual Try-On. ECCV 2022. Dataset access terms. github.com/aimagelab/dress-code
  4. [6]Virtual Try-On for Cultural Clothing: A Benchmarking Study (BD-VITON). arXiv:2603.07291, March 2026. Paper licensed CC BY-NC-SA 4.0; no separate dataset licence stated. arxiv.org/abs/2603.07291
  5. [14]Google Cloud. Virtual Try-On on Vertex AI (virtual-try-on-001). Documentation. docs.cloud.google.com/vertex-ai/generative-ai/docs/models/imagen/virtual-try-on-preview-08-04
  6. [15]JD.com. Oxygen-TryOn. arXiv:2607.21694, July 2026. arxiv.org/abs/2607.21694
  7. [16]Feng, Chen, Shan and Kemelmacher-Shlizerman. Layering Virtual Try-On. ECCV 2026. arXiv:2607.22924. arxiv.org/abs/2607.22924
  8. [17]Zhou et al. Learning Flow Fields in Attention for Controllable Person Image Generation (Leffa). CVPR 2025. arXiv:2412.08486. arxiv.org/abs/2412.08486
  9. [18]Chong et al. CatVTON: Concatenation Is All You Need for Virtual Try-On with Diffusion Models. ICLR 2025. arXiv:2407.15886. arxiv.org/abs/2407.15886
  10. [19]Rajput and Aneja. IndoFashion: Apparel Classification for Indian Ethnic Clothes. CVPR Workshops 2021. arXiv:2104.02830. arxiv.org/abs/2104.02830
  11. [25]Xu et al. OOTDiffusion: Outfitting Fusion Based Latent Diffusion for Controllable Virtual Try-On. AAAI 2025. arXiv:2403.01779. arxiv.org/abs/2403.01779
  12. [26]Choi et al. Improving Diffusion Models for Authentic Virtual Try-On in the Wild (IDM-VTON). ECCV 2024. arXiv:2403.05139. arxiv.org/abs/2403.05139
  13. [27]Jiang et al. FitDiT: Advancing the Authentic Garment Details for High-Fidelity Virtual Try-On. arXiv:2411.10499. arxiv.org/abs/2411.10499
  14. [28]Qwen Team. Qwen-Image-2.1 licence (Qwen Research License Agreement), September 2026. github.com/QwenLM/Qwen-Image-2.1/blob/main/LICENSE
  15. [29]Black Forest Labs. FLUX.2 [klein] 9B licence (FLUX Non-Commercial License). huggingface.co/black-forest-labs/FLUX.2-klein-9B/blob/main/LICENSE.md
  16. [30]FASHN AI. FASHN VTON v1.5: Efficient Maskless Virtual Try-On in Pixel Space. Research page. fashn.ai/research/vton-1-5
  17. [31]Zhu et al. TryOnDiffusion: A Tale of Two UNets. CVPR 2023. arXiv:2306.08276. arxiv.org/abs/2306.08276
  18. [32]Morelli et al. LaDI-VTON: Latent Diffusion Textual-Inversion Enhanced Virtual Try-On. ACM MM 2023. arXiv:2305.13501. arxiv.org/abs/2305.13501
  19. [33]Kim et al. StableVITON: Learning Semantic Correspondence with Latent Diffusion Model for Virtual Try-On. CVPR 2024. arXiv:2312.01725. arxiv.org/abs/2312.01725
  20. [34]OmniTry: mask-free virtual try-on for wearable objects. arXiv:2508.13632. arxiv.org/abs/2508.13632
  21. [35]Voost: a diffusion transformer for joint virtual try-on and try-off. SIGGRAPH Asia 2025. arXiv:2508.04825. arxiv.org/abs/2508.04825
  22. [36]TAMF-VTON. arXiv:2607.14807, July 2026. arxiv.org/abs/2607.14807
  23. [37]RealFit. arXiv:2609.25881, September 2026. arxiv.org/abs/2609.25881
  24. [38]FastFit: cacheable multi-reference virtual try-on. arXiv:2508.20586. arxiv.org/abs/2508.20586
  25. [39]CORAL: correspondence-aligned virtual try-on. arXiv:2602.17636, 2026. arxiv.org/abs/2602.17636
  26. [40]FASHN AI. fashn-vton-1.5 model card, Hugging Face. huggingface.co/fashn-ai/fashn-vton-1.5
  27. [41]VTEdit-Bench. ECCV 2026. arXiv:2603.11734. arxiv.org/abs/2603.11734
  28. [42]TryOnReward. arXiv:2609.13259, September 2026. arxiv.org/abs/2609.13259
  29. [43]DirectTryOn. arXiv:2605.12939, May 2026. arxiv.org/abs/2605.12939
  30. [44]FASHN AI. Open-sourcing FASHN VTON v1.5. Blog post, 2026. fashn.ai/blog/fashn-vton-1-5-open-source-release
  31. [45]FASHN AI. API pricing. Help centre. help.fashn.ai/plans-and-pricing/api-pricing
  32. [46]ByteDance Seed. Seedream 4.0 officially released. September 2025. seed.bytedance.com/en/blog/seedream-4-0-officially-released-beyond-drawing-into-imagination
  33. [47]Google Cloud. Virtual try-on technology with Google Cloud AI (Meesho case). July 2024. cloud.google.com/blog/products/ai-machine-learning/virtual-try-on-technology-with-google-cloud-ai
  34. [48]Qwen Team. Qwen-Image-Edit-2511 model card (Apache-2.0). huggingface.co/Qwen/Qwen-Image-Edit-2511
  35. [49]Black Forest Labs. FLUX.2 [klein] base 4B model card (Apache-2.0). huggingface.co/black-forest-labs/FLUX.2-klein-base-4B
  36. [50]Meituan LongCat. LongCat-Image-Edit model card. huggingface.co/meituan-longcat/LongCat-Image-Edit
  37. [51]HiDream. HiDream-O1-Image model card (MIT). huggingface.co/HiDream-ai/HiDream-O1-Image
  38. [52]Garments2Look. CVPR 2026. arXiv:2603.14153. arxiv.org/abs/2603.14153
  39. [53]OpenVTON-Bench. arXiv:2601.22725, January 2026. arxiv.org/abs/2601.22725
  40. [54]VTBench: a hierarchical virtual try-on benchmark. arXiv:2505.19571, May 2025. arxiv.org/abs/2505.19571
  41. [55]VTONQA. arXiv:2601.02945, January 2026. arxiv.org/abs/2601.02945
  42. [56]VTON-IQA and VTON-QBench. arXiv:2603.13057, March 2026. arxiv.org/abs/2603.13057