Skip to content
Clothsy AI Talk to us

Chapter 6AThe programme6 min readVersion 1.2 · 30 September 2026

Research outputs

Four papers, a public benchmark and toolkit, commercially usable image and video models, and live try-on on FabricVTON's own model.

6.1Publications

Paper 1

Every Garment, Every Body: A Global Benchmark and Fairness Audit of Virtual Try-On

CVPR 2027 workshop; NeurIPS 2027 Evaluations and Datasets

Month 6 to 7
Paper 2

Any Garment, Any View: Multi-View Pseudo-Paired Adaptation of Image-Editing Models for Virtual Try-On

ACM Multimedia 2027, BMVC 2027 or WACV 2028

Month 10
Paper 3

Taught by Photos: Commercially Clean Synthetic Pairs for Garment-Faithful Video Try-On

CVPR 2028

Month 14
Paper 4

Live Try-On: Garment-Faithful Streaming Virtual Try-On on One GPU

ECCV 2028 or SIGGRAPH 2028

Month 17 to 18
Table 37 Planned publications
OutputWorking titleContributionTarget venueMonth
Paper 1Every Garment, Every Body: A Global Benchmark and Fairness Audit of Virtual Try-OnFirst licensed try-on benchmark spanning garment families from several regions, with garment attributes, photo conditions and fairness strata; audit of at least ten systems; judge calibration studyCVPR 2027 workshop; NeurIPS 2027 Evaluations and Datasets6 to 7
Paper 2Any Garment, Any View: Multi-View Pseudo-Paired Adaptation of Image-Editing Models for Virtual Try-OnData engine for garments beyond stitched studio wear; multi-view garment conditioning; distilled student modelACM Multimedia 2027, BMVC 2027 or WACV 202810
Paper 3Taught by Photos: Commercially Clean Synthetic Pairs for Garment-Faithful Video Try-OnVideo data engine in which the image model re-dresses real clips; multi-view garment memory; a print and logo stability measure for videoCVPR 202814
Paper 4Live Try-On: Garment-Faithful Streaming Virtual Try-On on One GPUCausal student with persistent garment memory and self-forcing training; a live-camera benchmark; cost per stream-minuteECCV 2028 or SIGGRAPH 202817 to 18
Technical report (optional)Cost-Aware Serving of Virtual Try-On on Commodity GPUsQuality, latency and cost trade-offs for photo and live try-on, with blind human evaluationarXiv; WACV applications track4 to 6

6.2Datasets and benchmark

  • Global try-on benchmark. At least 2,000 consented test pairs across at least eight garment families from at least five world regions, with garment attribute labels, garment-presentation and photo-condition variants, and skin-tone and body-shape strata. Released for research with a datasheet.
  • Human rating set. Pairwise preferences and attribute judgements from raters in several regions, released to support calibration of automatic judges.
  • Proprietary training corpus. At least 20,000 licensed and synthetic pairs with a full provenance ledger. Kept private because licences and consents cover training, not redistribution.
  • Video hold-out. About 300 consented clips in studio, phone and live-camera conditions, with people and garments kept out of training. Released for research with a datasheet.
  • Proprietary video corpus. At least 7,000 real clips and 20,000 verified synthetic video pairs with a full provenance ledger. Kept private for the same reason as the image corpus.

6.3Models

  • Research model. A LoRA fine-tune of Qwen-Image-Edit-2511 for try-on of any garment, used for Paper 2.
  • Production model. A distilled student on FLUX.2 klein base 4B, deployed in FabricVTON's service once it beats the current model on the benchmark.
  • Video try-on model. An offline model on Wan2.1 14B for garment-faithful video try-on, used for Paper 3 and as the live teacher.
  • Live try-on model. A 1.3B causal student served over WebRTC behind FabricVTON's existing live page, replacing the rented engine once it wins on the hold-out.
  • Baseline results. Published scores for at least ten existing systems on the benchmark.

6.4Software

  • Evaluation toolkit, released open source under Apache-2.0, with the garment-attribute, skin-tone drift and body-distortion measures.
  • Judge calibration scripts that measure agreement between automatic judges and human raters.
  • Speed and cost laboratory for measuring latency and cost per try-on on cloud GPUs.
  • Live evaluation harness that measures time to first frame, sustained frame rate, frame latency and drift on the live-camera stress clips.

6.5Key performance indicators

Table 38 Key performance indicators
IndicatorTargetEvidenceMonth
Benchmark sizeAt least 2,000 image test pairs across at least eight garment families from at least five world regionsDatasheet5
Fairness coverageSubjects balanced across all ten Monk Skin Tone groups and across body shapes, including plus sizesDatasheet5
Systems auditedAt least tenPaper 16
Rater agreementKrippendorff's alpha of at least 0.6 on garment attributesPilot report2
Image model qualityStatistically significant human-preference win over the best open baseline on garment fidelityPaper 29
No regressionNo significant loss on standard metrics against the current production modelEvaluation report9
Image latencyDistilled model at or below 13 seconds on an L40SSpeed laboratory report11
Video dataAt least 7,000 real clips and 20,000 verified synthetic pairs with provenanceData card11
Video model qualityBeats open video baselines on garment fidelity over time on the 300-clip hold-outPaper 313
Live speedAt least 15 fps at 512p on one GPU, with the first frame in under 2 secondsLive evaluation report16
Live costUnder USD 0.12 per stream-minute, a tenth of the rented engineServing report16
Live stabilityPrints and logos stable across 60-second sessions by text accuracy and garment similarityLive evaluation report16
PublicationsTwo image-track papers submitted by month 10, and two video-track papers by month 18Submission records18
Open releaseBenchmark, video hold-out and evaluation toolkit publicRepositories12 and 18

6.6Intellectual property and open science

Methods, benchmark results and the evaluation toolkit will be published openly, with every contributor credited as a co-author. The benchmark test set and the video hold-out will be released under a research licence with datasheets. The production model, the training corpus and the provenance ledger remain FabricVTON's property, because the underlying licences and consents cover training but not redistribution. Any release of research weights will be decided per licence and consent, and stated in Paper 2.

6.8Conclusion