Upgrade to Pro
— share decks privately, control downloads, hide ads and more …
Speaker Deck
Features
Speaker Deck
PRO
Sign in
Sign up for free
Search
Search
The Natural Language Decathlon: Multitask Learn...
Search
Scatter Lab Inc.
July 10, 2019
Research
940
0
Share
Embed
Copy iframe code
Copy JS code
Copy link
Start on current slide
The Natural Language Decathlon: Multitask Learning as Question Answering
Scatter Lab Inc.
July 10, 2019
More Decks by Scatter Lab Inc.
See All by Scatter Lab Inc.
zeta introduction
scatterlab
0
2k
SimCLR: A Simple Framework for Contrastive Learning of Visual Representations
scatterlab
0
4.5k
Adversarial Filters of Dataset Biases
scatterlab
0
2.3k
Sparse, Dense, and Attentional Representations for Text Retrieval
scatterlab
0
2.4k
Weight Poisoning Attacks on Pre-trained Models
scatterlab
0
2.2k
Approximate Nearest Neighbor Negative Contrastive Learning for Dense Text Retrieval
scatterlab
0
2.6k
Beyond Accuracy: Behavioral Testing of NLP Models with CheckList
scatterlab
0
2.4k
Open-Retrieval Conversational Question Answering
scatterlab
0
2.3k
What Can Neural Networks Reason About?
scatterlab
0
2.3k
Other Decks in Research
See All in Research
クラウド・AI 時代の研究開発 DX / R&D Digital Transformation
hariby
0
130
[最先端NLP勉強会2026] Agentic Rubrics as Contextual Verifiers for SWE Agents
rfujii
1
340
Pretrain Where? Investigating How Pretraining Data Diversity Impacts Geospatial Foundation Model Performance
satai
3
130
OWASP AISVS - C7
shiell
1
640
2026年度 生成AI を活用した論文執筆ガイド/ワークショップ / 2026 Academic Year Guide to Writing Papers Using Generative AI - Workshop
ks91
PRO
0
230
[最先端NLP勉強会2026] Checklists Are Better Than Reward Models For Aligning Language Models
nzw0301
1
310
SLAMはどこまで解決されたのか?
tomonom
0
1.3k
Language and AI
ayaniwa
0
220
Spatial Active Noise Control Based onSound Field Interpolation Incorporating Physical Constraints
skoyamalab
0
180
マーケットストリート 社会実験2024 in 秋葉原ジャンク通り 調査報告書
izumiyama_lab
1
140
SoftMatcha 2: 1兆語規模コーパスの超高速かつ柔らかい検索
e869120_sub
7
3.8k
HackSick vol.7 LT資料【LLMアーキテクチャ入門・事前学習時の躓き所解説】 スパースなAttention・状態空間モデル
rikkabotan7
0
170
Featured
See All Featured
GraphQLの誤解/rethinking-graphql
sonatard
75
12k
Embracing the Ebb and Flow
colly
88
5.2k
Bash Introduction
62gerente
615
220k
How to train your dragon (web standard)
notwaldorf
97
6.8k
Designing for humans not robots
tammielis
254
26k
Marketing Yourself as an Engineer | Alaka | Gurzu
gurzu
0
290
Discover your Explorer Soul
emna__ayadi
2
1.3k
Statistics for Hackers
jakevdp
799
230k
Lightning talk: Run Django tests with GitHub Actions
sabderemane
0
250
How to Get Subject Matter Experts Bought In and Actively Contributing to SEO & PR Initiatives.
livdayseo
0
190
How People are Using Generative and Agentic AI to Supercharge Their Products, Projects, Services and Value Streams Today
helenjbeal
1
310
10 Git Anti Patterns You Should be Aware of
lemiorhan
PRO
659
62k
Transcript
스캐터랩(ScatterLab) ੌ࢚ച ੋҕמ Scatterlab ML Technical Seminar Session 2 (QA):
백영민 The Natural Language Decathlon: Multitask Learning as Question Answering McCann et al. Salesforce Research Machine Learning Engineer
#1. Concept
!3 Multitask Learning ৈ۞ о taskܳ э ण೧ࠁ!
• Method: ౠ objectiveܳ оҊ णػ model/representationਸ ܲ downstream taskী
ਊೞח Ѫ • ੌ߈ਵ۽ Language Modeling١ Natural Language ߈ੋ ౠࢿਸ णೡ ࣻ ח objective ਊ • : • Random initializeীࢲ दೞח Ѫ ࠁ જ Ѿҗܳ ࠁҊ, ࡅܲ ࣻ۴ਸ оמೞѱ ೧ષ • ߑध • Representation: Word2Vec, Glove ١ fixed representation, ULMFit, ELMO ١ intermediate representation(context aware)ਸ downstream taskীࢲ ਊೞח Ѫ(߹ب model ઓ) • Model: BERT, GPT١ pre-trainingী ਊ೮؍ modelਸ downstream taskীࢲ “fine-tuning” !4 #1 Concept Transfer Learning
• Method: ৈ۞ taskٜਸ ೞա ݽ؛۽ زदী णदఃח Ѫ •
Chunking, POS tagging, NER, SRL, dependency parsing, NLI ١ NLP taskٜਸ زੌೠ ݽ؛۽ زदী ण • : • ৈ۞ taskܳ زदী modelingೡ ࣻ • ੜ णغݶ ౠ taskী ೠػ Ѫ ইצ “General Representation”ਸ ਸ ࣻ • Zero-shot Learning, Meta-Learning ١ ਊ оמࢿ • ୭Ӕ ഝߊ োҳо ܖযҊ ח ࠙ঠ • ই singletask learningী ࠺Үؼ݅ೠ ࢿמਸ ࠁৈҊ ঋ݅, challengingೠ োҳٜ ݆ ܖযҊ . • Image classification + NLP • ࣗѐೡ ֤ޙ “MQAN” singletask learningী Ӓա݃ ࠺Үؼ݅ೠ ࢿמਸ ࠁ(BERT ֤ޙ ੑפ) !5 #1 Concept Multitask Learning
#2. Method
!7 Approach ݽٚ taskܳ QAഋधী ݏࠁ!
• য questionী ೠ ਸ “contextղীࢲ” ח ޙઁܳ ಿ (
Context ղࠗী Ҋ о) • Ex) SQuAD, RACE… • ࠁా द index৬ indexܳ ח ߑध !8 #2 Approach Question Answering
!9 #2 Approach Question Answering Idea: ݽٚ taskٜਸ “QAഋध”ਵ۽ ٜ݅যࢲ
“Multitask Learning”ਸ ೧ࠁ! QA Translation Summary NLI Sentiment Analysis
• ୨ 10ѐ taskܳ Multitask Learning !10 #2 Approach Decathlon
!11 Model Architecture Multitask Question Answering Network(MQAN)
!12 #2 Architecture Overview
!13 #2 Architecture I/O • Input: • Q: Question Sentences
• C: Context Sentences • A: Answer Sentences(for generation - autoregressive) • Output: • General QA: Contextীࢲ द, indexܳ • Q(Question Sentence) + C(Context Sentences) + Outer Vocabulary(Generation) ী ࢶఖ
!14 #2 Architecture Feature - Input Representation
!15 #2 Architecture Feature - Alignment <dummy dataܳ ֍ח ਬ>
!16 #2 Architecture Feature - Dual Coattention
!17 #2 Architecture Feature - Compression & Self-Attention
!18 #2 Architecture Feature - Answer Representation
!19 #2 Architecture Feature - Answer Representation
!20 #2 Architecture Feature - Answer
!21 Training Strategy Curriculum learning
!22 #3 Traning Strategy Multitask Learning Strategy • Round-robin Algorithm
• CPU scheduling ߑߨ ೞա۽ ஹೊఠ ਗਸ ࢎਊೡ ࣻ ח ӝഥܳ “۽ࣁझٜীѱ ҕ”ೞѱ ࠗৈ • п ۽ࣁझী ੌदрਸ ೡ, ೡػ दр աݶ Ӓ ۽ࣁझ ਫ਼द ࠁܨ, ܲ ۽ࣁझীѱ ӝഥܳ ષ • п Taskী ੌߓܳ ೡ, ೡػ ߓо աݶ Ӓ Taskਸ ਫ਼द ࠁܨ, ܲ Taskীѱ ӝഥܳ ષ • Fully Joint • п Task ࣽࢲܳ ҊೞҊ, round-robin algorithmਸ ా೧ batchܳ sampling • Single-task trainingীࢲ iterationਵ۽ب ࣻ۴೮؍ taskٜ ੜ غ݅ աݠח single-task݅ ण೮ ਸ ٸ݅ఀ ࢿמਸ ࠁৈ ޅೣ
!23 #3 Traning Strategy Multitask Learning Strategy • Curriculum Strategy
• Curriculumਸ ٜ݅যࢲ learningदெࠁ! • First Phase:࠺Ү ए taskٜਸ ݢ ण -> Second Phase:য۰ taskٜਸ ण • First Phase(SST, QA-SRL, QA-ZRE, WOZ, WikiSQL, MWSC) -> Second Phase(Others)
!24 #3 Traning Strategy Multitask Learning Strategy • Curriculum Strategy
• Curriculumਸ ٜ݅যࢲ learningदெࠁ! • First Phase:࠺Ү ए taskٜਸ ݢ ण -> Second Phase:য۰ taskٜਸ ण • First Phase(SST, QA-SRL, QA-ZRE, WOZ, WikiSQL, MWSC) -> Second Phase(Others) • Anti-Curriculum Strategy • Curriculumী “߈(Anti)ೞח” ۚ - Curriculum Learning Bengio et al. [2009] • ए taskח ܲ taskٜী بਸ ࣻ ח ਬਊೠ representationਸ णೡ ࣻ হ! • First Phase:য۰ taskٜਸ ݢ ण -> Second Phase:ए taskٜਸ ण • First Phase(SQuAD, IWSLT, CNN/DM, MNLI) -> Second Phase(Others)
#3. Result
!26 Result ־о־о ੜ೮ա?
!27 #1single vs multitask Single vs Multitask Training
!28 #2 curriculum Training Strategy
!29 #2 curriculum Pointer weight distribution
!30 Q & A хࢎפ