Most try-on research is judged one image at a time. Most shopping is not. A product page shows the garment from the front, the side and the back; people turn in front of a mirror; video is increasingly how clothes are shown. If a try-on is right in one pose and subtly different in the next, with the print shifted, the colour a little warmer and the hem at a different height, people stop trusting all of the images, including the good one.
Why generations drift
Diffusion models are stochastic, and try-on inherits that. Three things make consistency hard:
- Independent sampling. Each image starts from its own random noise. Small details, the ones that are not strongly pinned by the inputs, come out slightly different every time.
- Hidden regions. A product photo usually shows the front. When the person turns, the model has to invent the side seam, the back and the inside of a sleeve, and it can invent them differently in every view.
- Nothing ties the views together. In the simplest setup, each pose is its own independent request. Nothing in the model knows that the previous image exists.
Pose as a condition
Try-on models are told where the body is through pose and body representations computed from the person photo. Common choices include 2D keypoints from systems such as OpenPose [1], whole-body keypoints from DWPose [3], dense correspondence from DensePose, which maps image pixels to a 3D body surface [2], and human parsing to decide which regions are clothing [4].
These signals are essential, and they are good at what they do. But they describe the body, not the garment. They make sure the sleeve is on the arm in every pose; they do not make sure it is the same sleeve.
Multiple views and video
Research on consistency has approached it from two sides.
- Multi-view try-on. MV-VTON dresses a person seen from any viewpoint, using both front and back images of the garment, so that the back of the garment is taken from the product rather than invented [5].
- Video try-on. FW-GAN generated try-on video from a person, a garment and a pose sequence, using optical flow to warp earlier frames forward [6]. More recently, Tunnel Try-on focuses a diffusion model on a smoothed region that follows the clothing through the video, to keep detail and keep motion coherent [7].
Video makes the problem visible in a way still images do not. A small difference between two photos might go unnoticed; the same difference between consecutive frames is flicker.
Toward consistent garments
The direction we find most promising is to treat the garment as one thing, computed once and reused everywhere:
- A shared garment representation. Encode the product once and feed the same features to every view. Architectures with a dedicated garment branch, such as IDM-VTON’s garment UNet, already separate the garment from the person in a way that makes this natural [8].
- Condition views on each other. Let later views see earlier ones, so that anything the model had to invent is invented once.
- Measure sets, not samples. Compare garment regions across all views of one outfit, and measure flicker in video, alongside the usual single-image scores.
Consistency is one of our four open problems. It is also the one most tied to where shopping is going, toward more angles, more motion and more video.
References
- Cao et al. OpenPose: Realtime Multi-Person 2D Pose Estimation Using Part Affinity Fields. IEEE TPAMI, 2021.
- Güler, Neverova & Kokkinos DensePose: Dense Human Pose Estimation In The Wild. CVPR, 2018.
- Yang et al. Effective Whole-body Pose Estimation with Two-stages Distillation. ICCV Workshops (CV4Metaverse), 2023.
- Li et al. Self-Correction for Human Parsing. IEEE TPAMI, 2022.
- Wang et al. MV-VTON: Multi-View Virtual Try-On with Diffusion Models. AAAI, 2025.
- Dong et al. FW-GAN: Flow-navigated Warping GAN for Video Virtual Try-on. ICCV, 2019.
- Xu et al. Tunnel Try-on: Excavating Spatial-temporal Tunnels for High-quality Virtual Try-on in Videos. ACM Multimedia, 2024.
- Choi et al. Improving Diffusion Models for Authentic Virtual Try-on in the Wild. ECCV, 2024.
