Skip to content
Clothsy AI Talk to us

RESEARCH AREA 01

Visual Understanding

Models that understand people, objects, garments and environments.

Before a model can dress someone, it has to understand the photo it was given.

Every try-on starts with perception. The model needs to know where the person is, how they are standing, which pixels are skin, hair, background and clothing, and roughly what shape the body is underneath. It needs to understand the product photo too: what kind of garment it is, where the collar, sleeves and hem are, and which details matter.

These are classic computer-vision problems: human parsing, pose estimation, segmentation and dense body correspondence. Try-on stresses them in unusual ways. Loose, layered or dark clothing hides the body. Product photos arrive in every style, from flat lays and ghost mannequins to model shots. And errors compound: a parsing mistake at the start becomes a visible artifact at the end.

Much of this layer is built on strong open research, such as pose estimators, human parsers and promptable segmentation models. Our work is in making it reliable on real shopper photos and real catalogues, and in knowing when an input will not produce a good result.

Person photo (input)
Person photo
Product photo of a garment (input)
Product photo
Generated try-on result (output)
Generated try-on

Keep from the personFace and identity, pose, body shape, skin, hair, background.

Take from the productGarment shape, colour, fabric texture, print, logos, details like buttons and seams.

FigureThe task. A person photo and a product photo go in; a new image of that person wearing that product comes out. Images: the Clothsy AI demo set.

Questions we are exploring

  1. How do we estimate body shape reliably under loose or layered clothing?
  2. Can one model understand garment structure (collar, sleeves, closures, hem) from any style of product photo?
  3. Can we tell, before generating anything, that a photo will give a poor result, and tell the person why?

Further reading

  1. Cao et al. OpenPose: Realtime Multi-Person 2D Pose Estimation Using Part Affinity Fields. IEEE TPAMI, 2021.
  2. Güler, Neverova & Kokkinos DensePose: Dense Human Pose Estimation In The Wild. CVPR, 2018.
  3. Yang et al. Effective Whole-body Pose Estimation with Two-stages Distillation. ICCV Workshops (CV4Metaverse), 2023.
  4. Li et al. Self-Correction for Human Parsing. IEEE TPAMI, 2022.
  5. Ravi et al. SAM 2: Segment Anything in Images and Videos. ICLR, 2025.

JOIN THE RESEARCH TEAM

Work on the open problems with us.

Students, researchers and engineers. The application takes about two minutes.