Skip to content
Clothsy AI Talk to us

OPEN PROBLEMS

Hard, unsolved, and worth solving.

These are the problems our research team works on now. Each one is visible to a shopper in a single image, and none of them is solved by the field yet.

Virtual try-on looks like one task and behaves like four. A model has to carry every detail of the product across the move onto a body, make the fabric hang the way that fabric would, stay consistent from one pose to the next, and do all of it quickly and cheaply enough to run while someone is shopping. We treat each as a research problem: define it precisely, measure it on data the model has not seen, then work on it through experiments and ablations.

01 · Fabric fidelity through warping

A try-on has to move a flat product photo onto a body that is posed, turned and shaped differently. Every stripe, letter and logo has to survive the move. Small errors, like a bent letter, a smeared check or a pattern that changes scale, are exactly what a shopper notices first, and exactly what whole-image scores barely register.

Explicit warpTPSMove a control grid, warp the productimage, then blend it into the person.Implicit correspondenceEach output region looks up thegarment features it needs.
Figure 1Two ways to move a garment onto a body. Older systems warp the product image with an explicit transform, then blend it in. Diffusion systems learn the correspondence inside the network, through attention.

Why it is hard

  • Printed detail lives in a few pixels, so average image-quality metrics hardly move when it is wrong.
  • Latent diffusion models work in a compressed space, and fine print can be lost in the encoder before the generator ever sees it.
  • The garment must deform because it is on a body, but its print must not distort beyond what the fabric physically allows.

Directions we are exploring

  • Garment encoders that keep high-frequency detail alongside semantic features.
  • Losses and evaluation weighted toward printed and detailed regions.
  • Refinement passes for small, detailed areas of the image.

How we measure progress

  • Paired reconstruction on held-out garments, scored on garment crops.
  • Legibility checks for text and logos.
  • Side-by-side human preference focused on detail.

Read more: Evaluating texture fidelity in VTO

02 · Drape and folds

The same garment looks different on different bodies. Fabric weight, stiffness and stretch decide where cloth clings, where it falls free and how many folds form. A try-on that copies the product photo's own folds onto every body looks wrong, and it says the wrong thing about fit.

NarrowAverageBroad
Figure 2Drape. The same dress on three body shapes. Where the fabric touches the body, where it falls free, and how many folds form all change with the body underneath.

Why it is hard

  • A product photo shows how the garment falls on a mannequin or a table, not on this body.
  • Material properties are not labelled; they have to be inferred from appearance.
  • Physically based cloth simulation is accurate, but needs 3D garment patterns that product photos do not come with.

Directions we are exploring

  • Conditioning on body shape as well as pose.
  • Priors learned from physics-based cloth simulation.
  • Separating a garment's own folds from the folds a body creates.

How we measure progress

  • Test sets of the same garment on different bodies.
  • Human judgement of plausibility and of fit cues.
  • Where hems, waists and sleeves land relative to body landmarks.

Read more: Why fabric is harder than pixels

03 · Pose and camera consistency

People rarely judge a garment from one image. They look at several poses and angles, and increasingly at video. If a print shifts, a colour drifts or a hem jumps between frames, trust in every image drops.

PrintColourHemPrintColourHemPrintColourHem
Figure 3Consistency. The same garment on the same person in three poses: the print, the colour and the hem stripe must match in every frame, even when the arms and the camera move.

Why it is hard

  • Each generation is sampled independently, so small details vary from run to run.
  • Parts of the garment hidden in one pose, like the back or the inside of a sleeve, must be invented consistently in the next.
  • Video adds time: flicker is visible even when every frame looks good on its own.

Directions we are exploring

  • A shared garment representation reused across every view.
  • Multi-view and temporal conditioning.
  • Evaluation over sets of poses rather than single images.

How we measure progress

  • Similarity of garment regions across views.
  • Flicker and drift measures for video.
  • Human judgement of whole image sets.

Read more: Pose consistency in garment generation

04 · Fast, affordable inference

A try-on happens while someone is shopping. If it takes too long they leave, and if each image costs too much a store cannot offer it to every visitor. Speed is a research problem, not only an engineering one, because the fastest ways to sample a model change what it produces.

Original sampler: dozens of stepsDistilled student: a handful of stepsNoiseImage
Figure 4Why speed is a research problem. A diffusion model turns noise into an image over many denoising steps. Distillation trains a student that gets to a comparable image in a handful of steps, and the cost of a try-on falls with the step count.

Why it is hard

  • Diffusion models usually take many denoising steps, and each step is a full pass through a large network.
  • Few-step distillation can cost fidelity exactly where try-on needs it most: fine detail.
  • Serving has to hold quality under real traffic, with batching and reduced numerical precision.

Directions we are exploring

  • Few-step distillation tuned to preserve detail.
  • Smaller student models for serving.
  • Quantisation and batching without visible loss.

How we measure progress

  • Latency and cost per try-on at a fixed quality bar.
  • Fidelity before and after distillation, on the same test set.
  • Quality under realistic load.

Read more: Fast enough for a storefront

Want to work on these?We are building a research team of students, researchers and engineers. Apply in about two minutes or read about how the team works.

Back to research

JOIN THE RESEARCH TEAM

Work on the open problems with us.

Students, researchers and engineers. The application takes about two minutes.