Slide 1

Slide 1 text

ECCV2026採択 Reference-Free Image Quality Assessment for Virtual Try-On via Human Feedback Yuki Hirakawa1, 3, Takashi Wada1, Ryotaro Shimizu1, Takuya Furusawa1, Yuki Saito1, Ryosuke Araki2, Tianwei Chen1, Fan Mo1, and Yoshimitsu Aoki3 1 ZOZO Research, 2 ZOZO Inc., 3 Keio University

Slide 2

Slide 2 text

Virtual Try-On (VTON) ◼ VTON digitally simulate how a garment would look like on a given person. VTON • GAN • Diffusion ◼ A promising technology for bridging the gap between online fashion retail and physical stores. Aoki Media Sensing Lab. 1

Slide 3

Slide 3 text

VTON Still Frequently Generate Imperfect Results ◼ Displaying failed VTON results risks leading customers to make incorrect purchase decisions. The text printed on the shirt appears distorted. The neckline is depicted as a round neck instead of V-neck. VTON Aoki Media Sensing Lab. The stripe pattern on the sides is missing. 2

Slide 4

Slide 4 text

Reference-free Evaluation for VTON ◼ Score-based & Description-based evaluation of individual VTON images without ground-truth. Score-based Description-based The text printed on the shirt appears distorted. VTON The neckline is depicted as a round neck instead of V-neck. The stripe pattern on the sides is missing. Aoki Media Sensing Lab. 3

Slide 5

Slide 5 text

Reference-free Evaluation for VTON ◼ Score-based & Description-based evaluation of individual VTON images without ground-truth. Score-based Our Target Description-based The text printed on the shirt appears distorted. VTON The neckline is depicted as a round neck instead of V-neck. The stripe pattern on the sides is missing. Aoki Media Sensing Lab. 4

Slide 6

Slide 6 text

Key Requirements for Real-World Deployment ◼ No metrics enables automatic per-image evaluation without ground truth images. Reference-free Automation Per-image eval. SSIM × LPIPS × Human FID △ Ours × ×︎ ◼ Other desiderate o Lightweight and efficient o Minor errors should not be heavily penalized if the resulting VTON image looks natural. Aoki Media Sensing Lab. 5

Slide 7

Slide 7 text

VTON-QBench ◼ A large-scale human perception-aligned dataset for VTON quality evaluation. 1. Try-on image generation 2. Crowdsource 3-level evaluation 3. Dataset curation 789K Crowd workers Unreliable Reliable 431K 14 representative VTON models LLM(2) / DiT diff.(4) / U-Net diff.(5) / GAN(3) Aoki Media Sensing Lab. • • Dummy problems Repeatedly choosing the same option etc. 6

Slide 8

Slide 8 text

VTON-QBench ◼ Significant improvement of the questionnaire-level annotator agreement (Krippendorff's alpha) Discard ! Aoki Media Sensing Lab. 7

Slide 9

Slide 9 text

VTON-QBench ◼ A large-scale human perception-aligned dataset for VTON quality evaluation. 1. Try-on image generation 2. Crowdsource 3-level evaluation 3. Dataset curation 789K Crowd workers Unreliable Reliable 431K • • 14 representative VTON models LLM(2) / DiT diff.(4) / U-Net diff.(5) / GAN(3) Dummy problems Repeatedly choosing the same option etc. ◼ Dataset statistics (after curation) Aoki Media Sensing Lab. #Garment #VTON #Annotator #Annotation 13,153 62,688 13,838 431,800 8

Slide 10

Slide 10 text

Convert Discrete Annotation to Continuous MOS ◼ The final quality score S is computed as the average of numerical ratings from multiple annotators. Annotations to try-on query (G, P, V) ◼ The final quality score S provides an estimate of the average perceptual naturalness of VTON images. Aoki Media Sensing Lab. 9

Slide 11

Slide 11 text

VTON-IQA (Architecture) ◼ Interleaved Cross Attention (ICA) module is designed to measure: 1) Garment characteristic preservation; 2) Human pose, human identity, and background preservation. Quality Score in [-1. 1] Aoki Media Sensing Lab. 10

Slide 12

Slide 12 text

Generalization to Unseen Garment & Person ◼ SRCC with human judgements is evaluated on unseen garments and persons that were never seen during training. Aoki Media Sensing Lab. 11

Slide 13

Slide 13 text

Human vs. Automated Ranking of VTON Results Garment Human LPIPS SSIM VTON-IQA Aoki Media Sensing Lab. VTON 3 3 3 3 2 2 1 2 GT 1 1 2 Rank 1 12

Slide 14

Slide 14 text

Let’s Connect at Poster Session ! Session id OS2D-10 Paper Aoki Media Sensing Lab. Source Code & Dataset Demo 13