Upgrade to Pro
— share decks privately, control downloads, hide ads and more …
Speaker Deck
Sign up for free
Menu
Search
Features
All features
Private URLs
Password Protection
Custom URLS
Scheduled publishing
Remove Branding
Restrict embedding
Deck Collections
Notes
Features
All features
Private URLs
Password Protection
Custom URLS
Scheduled publishing
Remove Branding
Restrict embedding
Deck Collections
Notes
Explore
Featured decks
Featured speakers
Programming
Technology
Storyboards
Explore
Featured decks
Featured speakers
Programming
Technology
Storyboards
Pricing
Search
Sign in
Sign up for free
Adversarial Filters of Dataset Biases
Search
Scatter Lab Inc.
September 04, 2020
Research
2.3k
0
Share
Embed
Copy iframe code
Copy JS code
Copy link
Start on current slide
Adversarial Filters of Dataset Biases
Scatter Lab Inc.
September 04, 2020
More Decks by Scatter Lab Inc.
See All by Scatter Lab Inc.
zeta introduction
scatterlab
0
2.1k
SimCLR: A Simple Framework for Contrastive Learning of Visual Representations
scatterlab
0
4.6k
Sparse, Dense, and Attentional Representations for Text Retrieval
scatterlab
0
2.4k
Weight Poisoning Attacks on Pre-trained Models
scatterlab
0
2.2k
Approximate Nearest Neighbor Negative Contrastive Learning for Dense Text Retrieval
scatterlab
0
2.6k
Beyond Accuracy: Behavioral Testing of NLP Models with CheckList
scatterlab
0
2.4k
Open-Retrieval Conversational Question Answering
scatterlab
0
2.4k
What Can Neural Networks Reason About?
scatterlab
0
2.3k
Exploring the Limits of Transfer Learning with Unified Text-to-Text Transformer
scatterlab
0
2.3k
Other Decks in Research
See All in Research
医療LLMの現在地〜最新研究から社会実装までを考える〜
kento1109
1
1.8k
Source Code Diff Revolution
tsantalis
0
180
Research Engineerという仕事 / Research Engineering: Bridging Research and Business
chck
1
360
Fukui Shibiten 39 - AI Art
butchi
0
260
超効率化への挑戦:1bit LLMの現状と展望
yumaichikawa
0
780
進化?迷走?CasualConc ファミリーアプリの現在地 @ 英語コーパス学会 2026
casualconc
0
120
Easy to Guess, Hard to Verify: Lessons from AIMO 3 for Olympiad-Level AI Mathematics
corochann
0
120
[IR Reading 2026春 論文紹介] LLM-based Listwise Reranking under the Effect of Positional Bias (ECIR 2026) /IR-Reading-2026-Spring
koheishinden
PRO
0
450
人間中心の意思決定支援AI
yukinobaba
PRO
7
4.1k
20260624 NLP colloquium: 単一のhubテキストがCLIPを壊す:hubnessによる埋め込みの脆弱性特定
de9uch1
2
280
Sleuthcon Keynote - How Cybercriminals (ab)use AI
fr0gger
0
350
ふとした出会いで生まれたSkillが、 社内利用1位になるまで
mikimhk
22
24k
Featured
See All Featured
Documentation Writing (for coders)
carmenintech
77
5.6k
コードの90%をAIが書く世界で何が待っているのか / What awaits us in a world where 90% of the code is written by AI
rkaga
63
46k
Crafting Experiences
bethany
1
360
The untapped power of vector embeddings
frankvandijk
2
1.9k
WCS-LA-2024
lcolladotor
0
840
Design in an AI World
tapps
1
340
Evolution of real-time – Irina Nazarova, EuRuKo, 2024
irinanazarova
9
1.6k
Context Engineering - Making Every Token Count
addyosmani
9
1.2k
The Web Performance Landscape in 2024 [PerfNow 2024]
tammyeverts
12
1.3k
Everyday Curiosity
cassininazir
0
340
More Than Pixels: Becoming A User Experience Designer
marktimemedia
3
550
How To Speak Unicorn (iThemes Webinar)
marktimemedia
1
590
Transcript
Adversarial Filters of Dataset Biases ࢿࠁ (ML Research Scientist, Pingpong)
ݾର ݾର 1. োҳ ߓ҃ 2. AFLite 1. द: WinoGrande
ؘఠࣇ 2. ੌ߈ചػ ঌҊ્ܻ 3. प 1. Synthetic Data 2. NLP 3. Vision
োҳ ߓ҃ োҳ ߓ҃
‘߮݃ ؘఠࣇীࢲ ֫ ࢿמਸ ׳ࢿ೮Ҋ ೧ ޙઁܳ ೧Ѿ೮Ҋ ݈ೡ ࣻ
ਸө?’ • In-distribution పझࣇীࢲח ੜೞ݅ Out-of-distribution adversarial sampleীח ডೠ അ࢚ • Input-Output рী ب ঋ Spurious correlation ࢤ҂ӝ ٸޙ • ܳ ೧Ѿೠ ؘఠࣇਸ ٜ݅যঠ दझమਸ ઁ۽ ಣоೡ ࣻ োҳ ߓ҃ High Performance = Problem Solved?
োҳо domain-specificೠ spurious ಁఢਸ ࠙ܨ ߂ ೞҊ ܳ ઁѢೞח
ߑध • োҳ domain-specificೠ धҗ ҙী ઓ • ঌҊ્ܻ ࢸ҅о Ҋ۰ೞ ޅೠ biasח ழߡ ࠛо োҳ ߓ҃ Previous Approaches
AFLite AFLite
• ޙীࢲ ݺࢎо оܻఃח ࢚ਸ ݏח ޙઁ • SOTA ഛب
ড 90% → ݽ؛ Spurious correlationਸ ਊೞח ѱ ইקө? • (3), (4)ח ߃ հ݈ җ ҙ۲ ਸ ഛܫ ֫ই Word association݅ਵ۽ ޙઁܳ ಽ ࣻ AFLite Winograd Schema Challenge (WSC)
• ࢎۈ ؘఠࣇਸ ٜ݅ݶ ۠ Annotation artifactী ೠ Biasܳ
ೖೞӝ য۰ • AFLite۽ ఠ݂ೠ WinoGrande ؘఠࣇ ݽ؛ ഛبب ծҊ ܲ ߮݃۽ Transferب ੜؽ AFLite WinoGrande Dataset
1. ؘఠ ੌࠗ݅ਵ۽ RoBERTa fine-tune 2. Splitਸ ׳ܻ ೞݶࢲ RoBERTa
feature۽ linear classifier ण 3. Split పझࣇীࢲ ߬٬݅ਵ۽ ਸ औѱ ਸ ࣻ ח పझ ೞҊ ੋझఢझ߹۽ ঔ࢚࠶ ࣇী ୶о 4. ৈ۞ linear classifierо ਸ ݏ൦ ࠺ਯ Thresholdܳ ֈח Ѫ Top-kѐܳ ୭ઙ ؘఠࣇীࢲ ઁ৻ 5. ઁ৻غח ѐࣻо kѐо উ غѢա ਗೞח ӝ ؘఠࣇ ؼ ٸ ө 2~4 ߈ࠂ AFLite AFLite in WinoGrande
• ױয ӓࢿ݅ਵ۽ ಽ ࣻ ח ޙઁܳ Ѧ۞ն • ח
ష ۨ߰ Biasۄӝࠁח ҳઑੋ Ѫ۽ lexical-level heuristicਵ۽ח Ѧ۞ղӝ ൨ٝ AFLite Filtered Examples
• AFLiteܳ ৈ۞ بݫੋਵ۽ ഛೞҊ model-agnosticೞѱ ੌ߈ച • Contributions: 1.
࢚݅ intractableೠ AFOptܳ AFLite۽ Ӕࢎೡ ࣻ ਸ ࠁੋ. (Skip) 2. Vision, NLP ࠙ঠ ৈ۞ ؘఠࣇীࢲ प೧ AFLite ਬബࢿਸ ّ߉ஜೠ. 3. Biasܳ হঙ ؘఠࣇਵ۽ णೠ ݽ؛ ੌ߈ചо ੜؽਸ पਵ۽ ࠁੋ. 4. AFLite۽ ఠ݂ೞݶ ؊ بੋ ߮݃ ؘఠࣇਸ ٜ݅ ࣻ ਸ ࠁੋ. AFLite Adversarial Filters of Dataset Biases
: any feature extractor : a family of classification models
Φ M AFLite AFLite (Generalized)
Experiments Experiments
Biasing Dataset • Class-specificೠ ੋҕ featureܳ ؘఠ 75%ী ੑ, աݠח
random feature ੑ • Biased sample ੌࠗח ۨ࠶ ߄Է Results • Linear classifier۽ب ֫ ࢿמ ׳ࢿ • AFLiteܳ ਊೞݶ ࢚धੋ ࢿמਵ۽ جই১ Experiments Synthetic Data
• प ࢚: SNLI annotation artifactܳ ೖೠ out-of-distribution ؘఠࣇ 3ઙ
• Non-entailment ޙઁ ਬഋ߹۽ Zero-shot పझ Experiments NLP: Out-of-distribution Generalization
AFLite۽ ఠ݂ೠ ؘఠࣇ ݽٚ ݽ؛ীࢲ ࢿמ ѱ ڄয Experiments In-distribution
Benchmark Re-estimation: SNLI
Experiments In-distribution Benchmark Re-estimation: MultiNLI & QNLI
• : ImageNet ؘఠࣇ 20%۽ णೠ EfficientNet-B7 feature • ImageNet-A۽
ಣоೞפ AFLite-filtered ؘఠࣇਵ۽ ण೮ਸ ٸ ࢿמ ؊ જ Φ Experiments Vision: Adversarial Image Classification
ImageNet dev setਸ ఠ݂ೞҊ ಣо೮ਸ ٸ ࢿמ ೞۅ ؊ ఀ
Experiments In-distribution Image Classification
ӝઓীب ࠁҊػ ౠ ನૉী ೠ Bias, ݽনࠁ х݅ਵ۽ ҳ࠙ೞח ޙઁ
١җ Ѿਸ эೣ Experiments Filtered Examples
• Adversarial Filtering SWAG: A Large-Scale Adversarial Dataset for Grounded
Commonsense Inference [EMNLP’18] HellaSwag: Can a Machine Really Finish Your Sentence? [ACL’19] • AFLite WinoGrande: An Adversarial Winograd Schema Challenge at Scale [arXiv’19] Adversarial Filters of Dataset Biases [ICML’20] References References
хࢎפ✌ ୶о ޙ ژח ҾӘೠ ݶ ઁٚ ইې োۅ۽
োۅ ࣁਃ! ࢿࠁ (ML Research Scientist, Pingpong)
[email protected]