Upgrade to Pro
— share decks privately, control downloads, hide ads and more …
Speaker Deck
Sign up for free
Menu
Search
Features
All features
Private URLs
Password Protection
Custom URLS
Scheduled publishing
Remove Branding
Restrict embedding
Deck Collections
Notes
Features
All features
Private URLs
Password Protection
Custom URLS
Scheduled publishing
Remove Branding
Restrict embedding
Deck Collections
Notes
Explore
Featured decks
Featured speakers
Programming
Technology
Storyboards
Explore
Featured decks
Featured speakers
Programming
Technology
Storyboards
Pricing
Search
Sign in
Sign up for free
Tell-and-Answer: Towards Explainable Visual Que...
Search
onizuka laboratory
December 18, 2018
Research
94
0
Share
Embed
Copy iframe code
Copy JS code
Copy link
Start on current slide
Tell-and-Answer: Towards Explainable Visual Question Answering using Attributes and Captions
弊研究室で行なったEMNLP2018読み会の発表資料です。
onizuka laboratory
December 18, 2018
More Decks by onizuka laboratory
See All by onizuka laboratory
Phrase-Based & Neural Unsupervised Machine Translation
onilab
0
120
Card-660: A Reliable Evaluation Framework for Rare Word Representation Models
onilab
0
44
A Word-Complexity Lexicon and A Neural Readability Ranking Model for Lexical Simplification
onilab
0
160
Integrating Transformer and Paraphrase Rules for Sentence Simplification
onilab
0
70
An Auto-Encoder Matching Model for Learning Utterance-Level Semantic Dependency in Dialogue Generation
onilab
0
68
Generating More Interesting Responses in Neural Conversation Models with Distributional Constraints
onilab
0
110
Modeling Multi-turn Conversation with Deep Utterance Aggregation
onilab
0
100
Learning Semantic Sentence Embeddings using Pair-wise Discriminator
onilab
0
130
SGM: Sequence Generation Model for Multi-Label Classification
onilab
0
90
Other Decks in Research
See All in Research
GLIM とMegaParticles:正規分布近似の限界とタイトカップリング&パーティクルフィルタの進展 / GLIM and MegaParticles : Progress of the distribution representation in SLAM
koide3
0
820
MIRU2026 チュートリアル講演2:三次元データ処理の動向
nnchiba
6
4.8k
ros2-perf-multihost: 分散システムにおける客観的なアーキテクチャ評価フレームワーク
takasehideki
0
250
LA-Bench 2025:実験指示から実行可能手順を生成するためのデータセット/LA-Bench 2025: A Dataset for Generating Executable Experimental Procedures from Experimental Instructions
stktu
0
150
研究室単位での自律的 IPv6接続性確立に向けたAS共同運用モデルの提案と実証
reokashiwa
PRO
0
210
HAKARI-Bench - 実運用視点での情報検索モデル評価ベンチマーク
hotchpotch
1
740
最先端NLP勉強会2026 論文紹介:Reasoning with Sampling: Your Base Model is Smarter Than You Think (ICLR 2026 paper)
kogoro
4
600
NLP colloquium: AI Safety Survey
kanekomasahiro
2
1.1k
CVPR2026論文紹介_VLMにとって良いvision encoderとは何か?Rethinking Model Selection in VLM Through the Lens of Gromov-Wasserstein Distance
kobayashi31
1
210
Visual SLAM未来予測 / Future Prediction in Visual SLAM
koide3
1
1k
人間中心の意思決定支援AI
yukinobaba
PRO
7
4k
クラウド・AI 時代の研究開発 DX / R&D Digital Transformation
hariby
0
140
Featured
See All Featured
Large-scale JavaScript Application Architecture
addyosmani
515
110k
Fashionably flexible responsive web design (full day workshop)
malarkey
408
67k
エンジニアに許された特別な時間の終わり
watany
108
250k
I Don’t Have Time: Getting Over the Fear to Launch Your Podcast
jcasabona
35
2.8k
How to Ace a Technical Interview
jacobian
281
24k
Six Lessons from altMBA
skipperchong
29
4.5k
How to Build an AI Search Optimization Roadmap - Criteria and Steps to Take #SEOIRL
aleyda
1
2.2k
The untapped power of vector embeddings
frankvandijk
2
1.9k
Imperfection Machines: The Place of Print at Facebook
scottboms
270
14k
BBQ
matthewcrist
89
10k
Fight the Zombie Pattern Library - RWD Summit 2016
marcelosomers
234
17k
YesSQL, Process and Tooling at Scale
rocio
174
15k
Transcript
Tell-and-Answer: Towards Explainable Visual Question Answering using Attributes
and Captions Q. Li, J. Fu, D. Yu et al. EMNLP 2018 20181218
VQA CNN RNN ;* end-to-end !" → <2AE+H
C0 A9#5 2 4G' C0 <2I!".(AE)= >'1 %FI6?B,?:$D/-8 3@ VQA end-to-end &7 2 4G' <2AE)= 1
Visual Q&A Q: where is
the man swinging the racket? A: tennis court 2
Visual Q&A Q: what kind
of drink is in the glass? A: water 3
Visual Q&A Q: what is
walking next to the bus? A: cow 4
Visual Q&A Q: does the
man need a haircut? A: yes 5
Tell-and-Answer: Towards Explainable Visual Question Answering using Attributes and
Captions 6
"# where is the man swinging
the racket? yes no water tennis court ⋮ CNN RNN $ ! 7
7# 9+ D >(. <& D (CB D *
"6* D -2,4 end-to-end 2 3A $% @8181;?':/4-2 0),4 =C !581;?B 8
Tell-and-Answer: Towards Explainable Visual Question Answering using Attributes and
Captions 9
( ,# &+ Ø#'! Ø#(" ) ,( %*&$ ! Ø'
ver. Ø( ver. Ø ver. ( ) 10
11
@,→H3L0 !%$ %)=1M ;F$ % @, I87 H3L0 ResNet152
'#(' .? ØBAK6 ED "&'&( .?9 /$ % -NC> .? G+ ;F$ % cos N*2 J17 H3L0 cos N*2.? 5:4< H3)= 12
/ →3*) / ResNet152 LSTM 1 .%0
1', (e.g. BLEU) 4") .%) 2$5! cos 6# (+&- / 3*).% 13
7*='?8.-?9(-→64 #2(> 8.-9(- LSTM !" Ø LSTM &7!!)<%/
;5 $5') 3: softmax ,0+1 64#2 14
Tell-and-Answer: Towards Explainable Visual Question Answering using Attributes and
Captions 15
-VQA-real Ø1 & 3 ,1 10
)', *!#-min(#)*+,-. /010-/ 2),2 ,-.345 6 , 1) Ø%)' 10 " 3 ( $) + * 16
VQA vs.
17
vs.
18
VQA *!-# +$ .% -#"0,)( 0'/ Ø& -#" Ø
-#" ØNULL 1 -# $ 19
* '" Ø+! & )%
Ø+! & )% Ø+! & )% Ø+! & )% $ ,# - +! ( VQA & 20
tennis, ball, man, racket, hit, court, play, player,
swing, hold a man holding a tennis racket on a tennis court. tennis court & Q: where is the man swinging the racket? A: tennis court 21
bicycle, man, sit, eat, bike, look, outside, food,
person, table a man sitting at a table with a plate of food. beer & Q: what kind of drink is in the glass? A: water 22
street, bus, cow, city, walk, car, drive, stand,
road, white a cow that is walking in the street. car & Q: what is walking next to the bus? A: cow 23
woman, bear, teddy, hold, sit, glass, animal, large,
lady a woman holding a sandwich in her hands. yes & Q: does the man need a haircut? A: yes 24
30%
65% yes/no 80% 25
VQA 26
Tell-and-Answer: Towards Explainable Visual Question Answering using Attributes and
Captions 27
>1=9 5 2 3B# VQA #2 Ø7!6*A-"?&.% / Ø7!
<+; 0',4:(C $ 8 VQA =@ ) = 28