categories: 1,100+ • clothes • shoes • bags • books ImageNet (ILSVRC) • Number of categories: 1,000 • milk can • rugby ball • paper towel • earthstar • envelope • miniskirt, mini • cowboy hat, ten-gallon hat • Trolleybus, trolley coach • etc. • cars • handmade • cosmetics • toys • etc. The level of difficulty in distinguishing items is about the same for Mercari and ImageNet → Perhaps we can use machine learning to demonstrate the human ability to recognize items sold on Mercari?
is high even without data cleansing or parameter tuning, but is it as good as a human? Items are properly recognized for most categories Can we improve the recognition quality even more? What is happening here? Can we improve?
100cm~ > Jackets The numbers in parentheses indicate the level (score) of recognition accuracy and the sum of all cases being monitored is 1.0 • Women > Jackets/Outerwear > Pea coats (0.4849) • Men > Jackets/Outerwear > Pea coats (0.3372) • Babies/Kids > Boys outfit 100cm~ > Coats (0.0305) • Babies/Kids > Boys/girls outfit 100cm~ > Coats (0.0289) • Women > Jackets/Outerwear > Trench coats (0.0176) (Answer: the category selected by our customers)
• Sometimes multiple categories fit the image • The pink jacket with polka dots could be for boys as well as girls • Even humans have a hard time distinguishing items of slightly different sizes • Applies to clothes and other items like shoes Interestingly, the error tendencies are very similar to human mistakes.
> Toys > Musical box The ability to recognize “Looping” and “Mary” toys was premature; machine learning and data collection was insufficient for the diversity of these images. Can calculate item similarity levels using CNN’s middle layer Babies/Kids > Toys > Musical box (0.9111)
a child how to see, especially in the early years. They learn this through real-world experiences and examples. If you consider a child's eyes as a pair of biological cameras, they take one picture about every 200 milliseconds, the average time an eye movement is made. So by age three, a child would have seen hundreds of millions of pictures of the real world. That's a lot of training examples. So instead of focusing solely on better and better algorithms, my insight was to give the algorithms the kind of training data that a child was given through experiences, in both quantity and quality. “
In terms of category recognition of item images on Mercari • (While we haven’t set specific numeric goals) the technology has not surpassed human abilities • It is possible to demonstrate human’s recognition capabilities using existing recognition technology • There is a need to study greater-scale data sets to equip for diversity • Beyond item category, we can use image recognition technology to predict: • Brand • Item title • Item condition • Price range ...for even better customer experience
CVPR 2016. • Visualization technique called Class Activation Mapping • Applicable for general Convolutional Neural Networks • Also applicable for Inception-v3 What Info does Deep Neural Networks Use for Recognition?
tell: A neural image caption generator, CVPR 2015. • https://github.com/tensorflow/models/tree /master/im2txt • https://research.googleblog.com/2016/09/ show-and-tell-image-captioning-open.html Besides image categorization, research on generating accurate item descriptions is in progress. → Can we generate item titles from item images?
such as Deep Neural Networks and Deep Learning Babies/Kids > Kids shoes > Sandals Crocs Crocband Crocs ¥1,500 - ¥2,500 Pink Used over a total of 50 million pieces of item data for learning (We are currently at the accuracy testing stage, and quantitative test results are not ready for reporting.) Prediction of Item Details Including Item Title
can seem as though the prediction has surpassed human capabilities Prediction of Item Details Including Item Title Ralph Lauren polo shirt Men > Tops > Polo Shirt Louis Vuitton Monogram Hock Women > Accessories > Billfold wallet Bvlgari Pour Homme “Bvlgari” (written in Japanese) Cosmetics/Perfume/Beauty > Perfume > cologne (men)
no specific product or brand knowledge, it can seem as though the prediction has surpassed human capabilities Combi baby wipe warmer Babies/kids > Diaper/Toilet/Bath > Diapers “Mappuru” Seoul mini Entertainment/Hobby > Books > Maps/travel guides TV remote controller Electronics/smartphones/cameras > TV/video equipment > Others
was accurately generated from image features without OCR Has properly collected information that the item is a SHARP AQUOS remote controller TV remote controller Electronics/smartphones/cameras > TV/video equipment > Others
title was accurately generated from image features without OCR Has recognized that the item is a magazine about Seoul “Mappuru” Seoul mini Entertainment/Hobby > Books > Maps/travel guides
widely recognized by families/households with newborn babies and toddlers Has not recognized that it is a baby wipe warmer, or a baby good Combi baby wipe warmer Babies/kids > Diaper/Toilet/Bath > Diapers
have achieved human-level precision • Studied categorization modeling using Mercari item images • Our item recognition rates were far from ImageNet recognition rates • Categories with high rate of recognition errors: • Categories like “others” whose definition is not clear or specific • Categories defined by different sizes and gender (men’s, women’s, kids’, etc) • There is space for improvement • High parameter adjustment at the time of learning • Add machine learning data (especially for categories that consist of a variety of items) • Future development • Prediction to include not just categories but item titles, brand, price, color, etc • A prototype for predictions using above factors is already made; we are currently testing its accuracy • When there is a lack of product and brand knowledge, it seems as though the prediction has surpassed human capabilities • e.g. Combi baby wipe warmer / BVLGARI pour homme