- Virtual try-on (VTON)
- Generating a photograph of a person wearing a garment from separate photographs of the person and the garment
- Diffusion model
- A generative model that creates an image by removing noise step by step
- MMDiT
- Multimodal diffusion transformer, which processes image and condition tokens in shared attention
- LoRA
- Low-rank adaptation, a way to fine-tune a large model by training small added matrices
- Try-off
- Generating a flat product image of a garment from a photograph of someone wearing it
- Distillation
- Training a smaller or faster model to reproduce a larger model's outputs
- FID and KID
- Fréchet and Kernel Inception Distances, which measure how close generated images are to real ones as a set
- SSIM and LPIPS
- Structural and learned perceptual similarity between an output and its ground truth
- Kendall τ
- A rank correlation used here to measure agreement between metrics and human raters
- Krippendorff's alpha
- A measure of agreement between human annotators
- Monk Skin Tone scale
- A ten-point skin tone scale created by Dr Ellis Monk and published by Google for inclusive evaluation
- CIEDE2000
- A standard formula for perceived colour difference
- Garment family
- A group of garments that fail in similar ways, such as tops, outerwear or draped garments
- Garment presentation
- How the garment photograph shows the item: flat, folded as sold, on a hanger or mannequin, or worn by another person
- Draped or wrapped garment
- A garment whose worn shape comes from wrapping, folding or tying on the body rather than from its cut alone, such as a saree, sarong, shawl or wrap dress
- DPIIT
- Department for Promotion of Industry and Internal Trade, which recognises startups in India
- Video try-on
- Dressing a person in a garment across a whole video, keeping identity, motion and the garment consistent from frame to frame
- Live try-on
- Video try-on on a camera stream as it arrives, with each frame drawn in about 33 to 66 milliseconds
- Frames per second (fps)
- How many new images a video system produces each second; live video needs about 15 to 30
- Causal generation
- Drawing each frame only from past frames and the current input, never future ones, which live video requires
- Self-forcing
- Training a video model on its own generated frames so that small errors do not build up over time
- KV cache
- Stored attention keys and values from earlier frames, reused so that each new frame is fast to compute
- WebRTC
- The open standard browsers use for live video calls, used here to stream the camera to the GPU and back
- VFID
- Video Fréchet Inception Distance, a measure of how close generated videos are to real ones as a set
AppendixReference3 min readVersion 1.2 · 30 September 2026
Glossary
Plain-language definitions of the terms used across the programme.
