Generative AI

Virtual Try-On Diffusion Model for Wearit

Architected Wearit's core generative AI pipeline for virtual clothing try-on, from R&D to real-time production: a diffusion model built from scratch, trained on the Jean Zay supercomputer on 3M+ labeled fashion images, serving results in under 2 seconds.

Virtual Try-On Diffusion Model for Wearit
  • 3M+ labeled images
  • < 2s inference
  • Jean Zay / GENCI

The problem

Wearit builds virtual try-on for fashion: a shopper sees a garment on a person instead of on a flat product shot. The idea is simple to state and hard to execute. The output has to keep the garment exactly as sold, with its cut, color, print, logo and fabric texture intact, while fitting it naturally to a different body, pose and lighting. It also has to look like a photograph rather than a collage, and it has to be fast enough to sit inside a shopping experience.

Off-the-shelf image generation models do not solve this. General-purpose diffusion models are trained to produce plausible images, not to reproduce a specific product faithfully. A try-on result that changes the pattern on a shirt or invents a pocket is worse than no result, because it misrepresents what the customer is buying.

Our approach

We architected the core generative AI pipeline for virtual clothing try-on, from research through to real-time production.

A diffusion model built from scratch. Rather than adapting a general text-to-image checkpoint, we built a dedicated virtual try-on diffusion model. Diffusion-based try-on conditions generation on the garment image and on a representation of the person, so the model learns to transfer the garment onto the body while preserving its details. Owning the architecture end to end means every design choice, from how the garment is encoded to how the person is conditioned, serves the try-on task rather than general image synthesis.

Distributed training on a national supercomputer. The model was trained in distributed multi-GPU runs on A100 and H100 GPUs on Jean Zay, the French national supercomputer, under a GENCI compute grant. Training a generative model from scratch at this scale is a systems problem as much as a modeling one: data loading, checkpointing, and multi-node scaling all have to work reliably inside a shared HPC scheduler.

Data engineering at scale. The training set comprised more than 3 million custom-labeled fashion images. For try-on, the labels are what teach the model the structure of the problem: which pixels belong to the garment, which to the body, and how the two relate. Building that pipeline was a core part of the work, not a preprocessing step.

Production inference under 2 seconds. Diffusion models are iterative by nature, and a naive sampler is far too slow for a shopping flow. The production system returns a try-on result in under 2 seconds.

This combination of frontier modeling and production engineering is what our computer vision and research & development practices are built for. We describe the general techniques in more depth in our guide on how to build and train a virtual try-on diffusion model.

Results

  • 3M+ labeled images: a custom-labeled fashion dataset of more than 3 million images, built for the try-on task.
  • Under 2 seconds: production inference latency for a generated try-on image.
  • Jean Zay / GENCI: distributed multi-GPU training on A100 and H100 GPUs on the French national supercomputer, under a GENCI compute grant.

The latency figure is what turns a research model into a product. A try-on that takes half a minute is a demo; one that returns in under 2 seconds can be part of browsing a catalog.

What it takes to deploy

Shipping generative try-on is more than training a good model. A production deployment typically involves:

  • Garment fidelity checks. Evaluation has to catch changed prints, colors and logos, not just score overall image realism.
  • Coverage of the real catalog. Performance must hold across garment categories, body types, poses and photography styles that the business actually sells and receives.
  • Latency engineering. Sampling steps, model size and GPU serving have to be tuned together to hit an interactive budget at the expected traffic.
  • Compute planning. Training from scratch needs sustained access to large GPU clusters, whether through public HPC programs like GENCI or private infrastructure.

Where this applies

The same pattern, a task-specific diffusion model trained on carefully labeled data and engineered for real-time inference, applies wherever generated images must stay faithful to a real object: fashion and retail imagery, product visualization, and other conditional generation problems. Teams training large models on national or private GPU clusters face the same systems challenges, which is the focus of our work in R&D & frontier compute.

Building a generative model that has to be both faithful and fast? Get in touch.

Capabilities & industries

More projects