Lock in $30 Savings on PRO—Offer Ends Soon! ⏳
Speaker Deck
Features
Speaker Deck
PRO
Sign in
Sign up for free
Search
Search
Tell-and-Answer: Towards Explainable Visual Que...
Search
onizuka laboratory
December 18, 2018
Research
0
72
Tell-and-Answer: Towards Explainable Visual Question Answering using Attributes and Captions
弊研究室で行なったEMNLP2018読み会の発表資料です。
onizuka laboratory
December 18, 2018
Tweet
Share
More Decks by onizuka laboratory
See All by onizuka laboratory
Phrase-Based & Neural Unsupervised Machine Translation
onilab
0
120
Card-660: A Reliable Evaluation Framework for Rare Word Representation Models
onilab
0
36
A Word-Complexity Lexicon and A Neural Readability Ranking Model for Lexical Simplification
onilab
0
130
Integrating Transformer and Paraphrase Rules for Sentence Simplification
onilab
0
61
An Auto-Encoder Matching Model for Learning Utterance-Level Semantic Dependency in Dialogue Generation
onilab
0
57
Generating More Interesting Responses in Neural Conversation Models with Distributional Constraints
onilab
0
100
Modeling Multi-turn Conversation with Deep Utterance Aggregation
onilab
0
98
Learning Semantic Sentence Embeddings using Pair-wise Discriminator
onilab
0
120
SGM: Sequence Generation Model for Multi-Label Classification
onilab
0
80
Other Decks in Research
See All in Research
J-RAGBench: 日本語RAGにおける Generator評価ベンチマークの構築
koki_itai
0
1.1k
若手研究者が国際会議(例えばIROS)でワークショップを企画するメリットと成功法!
tanichu
0
120
HoliTracer:Holistic Vectorization of Geographic Objects from Large-Size Remote Sensing Imagery
satai
3
320
Thirty Years of Progress in Speech Synthesis: A Personal Perspective on the Past, Present, and Future
ktokuda
0
120
Open Gateway 5GC利用への期待と不安
stellarcraft
2
160
世界の人気アプリ100個を分析して見えたペイウォール設計の心得
akihiro_kokubo
PRO
63
34k
日本語新聞記事を用いた大規模言語モデルの暗記定量化 / LLMC2025
upura
0
360
"主観で終わらせない"定性データ活用 ― プロダクトディスカバリーを加速させるインサイトマネジメント / Utilizing qualitative data that "doesn't end with subjectivity" - Insight management that accelerates product discovery
kaminashi
15
15k
音声感情認識技術の進展と展望
nagase
0
400
MetaEarth: A Generative Foundation Model for Global-Scale Remote Sensing Image Generation
satai
4
490
IMC の細かすぎる話 2025
smly
2
780
論文紹介: ReGenesis: LLMs can Grow into Reasoning Generalists via Self-Improvement
hisaokatsumi
0
140
Featured
See All Featured
Optimizing for Happiness
mojombo
379
70k
Building a Modern Day E-commerce SEO Strategy
aleyda
45
8.3k
VelocityConf: Rendering Performance Case Studies
addyosmani
333
24k
Product Roadmaps are Hard
iamctodd
PRO
55
12k
RailsConf & Balkan Ruby 2019: The Past, Present, and Future of Rails at GitHub
eileencodes
141
34k
Building Adaptive Systems
keathley
44
2.9k
[RailsConf 2023] Rails as a piece of cake
palkan
58
6.2k
Responsive Adventures: Dirty Tricks From The Dark Corners of Front-End
smashingmag
254
22k
ReactJS: Keep Simple. Everything can be a component!
pedronauck
666
130k
Building a Scalable Design System with Sketch
lauravandoore
463
34k
Distributed Sagas: A Protocol for Coordinating Microservices
caitiem20
333
22k
Unsuck your backbone
ammeep
671
58k
Transcript
Tell-and-Answer: Towards Explainable Visual Question Answering using Attributes
and Captions Q. Li, J. Fu, D. Yu et al. EMNLP 2018 20181218
VQA CNN RNN ;* end-to-end !" → <2AE+H
C0 A9#5 2 4G' C0 <2I!".(AE)= >'1 %FI6?B,?:$D/-8 3@ VQA end-to-end &7 2 4G' <2AE)= 1
Visual Q&A Q: where is
the man swinging the racket? A: tennis court 2
Visual Q&A Q: what kind
of drink is in the glass? A: water 3
Visual Q&A Q: what is
walking next to the bus? A: cow 4
Visual Q&A Q: does the
man need a haircut? A: yes 5
Tell-and-Answer: Towards Explainable Visual Question Answering using Attributes and
Captions 6
"# where is the man swinging
the racket? yes no water tennis court ⋮ CNN RNN $ ! 7
7# 9+ D >(. <& D (CB D *
"6* D -2,4 end-to-end 2 3A $% @8181;?':/4-2 0),4 =C !581;?B 8
Tell-and-Answer: Towards Explainable Visual Question Answering using Attributes and
Captions 9
( ,# &+ Ø#'! Ø#(" ) ,( %*&$ ! Ø'
ver. Ø( ver. Ø ver. ( ) 10
11
@,→H3L0 !%$ %)=1M ;F$ % @, I87 H3L0 ResNet152
'#(' .? ØBAK6 ED "&'&( .?9 /$ % -NC> .? G+ ;F$ % cos N*2 J17 H3L0 cos N*2.? 5:4< H3)= 12
/ →3*) / ResNet152 LSTM 1 .%0
1', (e.g. BLEU) 4") .%) 2$5! cos 6# (+&- / 3*).% 13
7*='?8.-?9(-→64 #2(> 8.-9(- LSTM !" Ø LSTM &7!!)<%/
;5 $5') 3: softmax ,0+1 64#2 14
Tell-and-Answer: Towards Explainable Visual Question Answering using Attributes and
Captions 15
-VQA-real Ø1 & 3 ,1 10
)', *!#-min(#)*+,-. /010-/ 2),2 ,-.345 6 , 1) Ø%)' 10 " 3 ( $) + * 16
VQA vs.
17
vs.
18
VQA *!-# +$ .% -#"0,)( 0'/ Ø& -#" Ø
-#" ØNULL 1 -# $ 19
* '" Ø+! & )%
Ø+! & )% Ø+! & )% Ø+! & )% $ ,# - +! ( VQA & 20
tennis, ball, man, racket, hit, court, play, player,
swing, hold a man holding a tennis racket on a tennis court. tennis court & Q: where is the man swinging the racket? A: tennis court 21
bicycle, man, sit, eat, bike, look, outside, food,
person, table a man sitting at a table with a plate of food. beer & Q: what kind of drink is in the glass? A: water 22
street, bus, cow, city, walk, car, drive, stand,
road, white a cow that is walking in the street. car & Q: what is walking next to the bus? A: cow 23
woman, bear, teddy, hold, sit, glass, animal, large,
lady a woman holding a sandwich in her hands. yes & Q: does the man need a haircut? A: yes 24
30%
65% yes/no 80% 25
VQA 26
Tell-and-Answer: Towards Explainable Visual Question Answering using Attributes and
Captions 27
>1=9 5 2 3B# VQA #2 Ø7!6*A-"?&.% / Ø7!
<+; 0',4:(C $ 8 VQA =@ ) = 28