Upgrade to Pro
— share decks privately, control downloads, hide ads and more …
Speaker Deck
Sign up for free
Menu
Search
Features
All features
Private URLs
Password Protection
Custom URLS
Scheduled publishing
Remove Branding
Restrict embedding
Deck Collections
Notes
Features
All features
Private URLs
Password Protection
Custom URLS
Scheduled publishing
Remove Branding
Restrict embedding
Deck Collections
Notes
Explore
Featured decks
Featured speakers
Programming
Technology
Storyboards
Explore
Featured decks
Featured speakers
Programming
Technology
Storyboards
Pricing
Search
Sign in
Sign up for free
ICLR2017読み会@DeNA/iclr2017atDeNA_VLAE
Search
Masaki Kozuki
June 17, 2017
Research
29k
2
Share
Embed
Copy iframe code
Copy JS code
Copy link
Start on current slide
ICLR2017読み会@DeNA/iclr2017atDeNA_VLAE
Masaki Kozuki
June 17, 2017
More Decks by Masaki Kozuki
See All by Masaki Kozuki
Sanity Checks for Saliency Maps explained in Japanese language
crcrpar
0
2.7k
Deep Learning for clothes and changing pose
crcrpar
0
940
夏のトップカンファレンス論文読み会 / InnovationMeetup20170918csn_cvpr2k17
crcrpar
3
1.5k
iclr読み会 / iclrjp2017vlae
crcrpar
3
1.1k
Other Decks in Research
See All in Research
RS-Agent: Automating Remote Sensing Tasks through Intelligent Agent
satai
3
580
[IR Reading 2026春 論文紹介] LLM-based Listwise Reranking under the Effect of Positional Bias (ECIR 2026) /IR-Reading-2026-Spring
koheishinden
PRO
0
450
[WebDB2026]セレンディピティ指向推薦システム再考 ―セレンディピティの原義・類型・発生過程に基づく設計指針―
recsyslab
PRO
0
170
高性能計算機クラスタを用いた大規模点群処理による森林の単木抽出と構造解析
kentaitakura
1
140
CDCL を用いた MILP の厳密解法
imai448
0
280
実例から見るLLMのマンガ理解:実務VQAタスクによる長期的文脈と視覚情報の定性評価
kzmssk
0
160
完全自律LLMエージェントでDIVER OSINT CTF 2026に挑戦してみた
analokmaus
0
700
敵対生成プロンプト同時探索による内省型プロンプト最適化
kinoue_smarthr
0
430
論文紹介:Doc-to-LoRA: Learning to Instantly Internalize Contexts
yukako_nakano
0
180
長時間動画QAにおけるマルチエージェント推論 ・SVAgent: Storyline-Guided Long Video Understanding via Cross-Modal Multi-Agent Collaboration
murakawatakuya
1
210
【Zozo Research 技術共有会】三次元領域の現在と展望
mickey_0226
3
670
MM-OVSeg: Multimodal Optical–SAR Fusion for Open-Vocabulary Segmentation in Remote Sensing
satai
3
150
Featured
See All Featured
Believing is Seeing
oripsolob
1
230
AI: The stuff that nobody shows you
jnunemaker
PRO
10
1.1k
Chasing Engaging Ingredients in Design
codingconduct
0
340
How to Get Subject Matter Experts Bought In and Actively Contributing to SEO & PR Initiatives.
livdayseo
0
200
Navigating Algorithm Shifts & AI Overviews - #SMXNext
aleyda
1
1.6k
How to Talk to Developers About Accessibility
jct
2
560
ラッコキーワード サービス紹介資料
rakko
1
5.1M
Beyond borders and beyond the search box: How to win the global "messy middle" with AI-driven SEO
davidcarrasco
3
270
Winning Ecommerce Organic Search in an AI Era - #searchnstuff2025
aleyda
2
2.2k
Building Flexible Design Systems
yeseniaperezcruz
330
41k
Bash Introduction
62gerente
615
220k
Templates, Plugins, & Blocks: Oh My! Creating the theme that thinks of everything
marktimemedia
31
2.9k
Transcript
Variational Lossy Autoencoder ICLR 2017 ಡΈձ @ DeNA @crcrpar 2017/6/17
1 / 25
จ • Variational Lossy Autoencoder • Xi Chen (UC Berkeley,
OpenAI), Diederik P. Kingma (OpenAI), Tim Salimans (OpenAI), et al. • දݱֶशͰજࡏมΛ׆༻͢Δ • Bits Back Coding Ͱ VAE ͷજࡏมʹ͍ͭͯͷߟ • જࡏมΛ lossy ʹ͢Δ • જࡏม z ͷ p(z), q(z|x) Λॊೈʹ • decoder ʹ PixelCNN 2 / 25
දهʹ͍ͭͯ • x ∈ Rd: σʔλ. x = ( x0
. . . xd )⊤ • x<i : x ͷ index ͕ i ະຬͷશཁૉ ( x0 . . . xi−1 )⊤ • z: જࡏม • pdata (x): σʔλΛੜ͢Δਅͷ • DKL (p∥q): p ͷ q ʹର͢Δ Kullback Leibler Divergence • θ: ϞσϧʢNNʣͷύϥϝʔλ • AR: PixelCNN ͳͲͷࣗݾճؼܕ NN • H, H: Τϯτϩϐʔ 3 / 25
VAE తؔ log p(X) = ∑ N i=1 log p(x(i))
࣮ࡍͷతؔ L(x; θ) = Eq(z|x) [log p(x|z) − DKL (q(z|x)∥p(z))] - ਖ਼نԽͨ͠ autoencoder ͱΈΕΔɻ VAE ͷ՝ɾऑ • දݱྗ͕ߴ͗͢Δ decoder જࡏมΛແࢹ • જࡏม͕ͭใΛཧͰ͖ͳ͍ 4 / 25
1 ͳͥʁ ײతʹ ཧʢBits Back Codingʣ 2 VLAE ֓ཁ Autoregressive
Flow decoder: PixelCNN 3 ࣮ݧɾ݁Ռ Lossy Comprssion Density Estimation 5 / 25
1 ͳͥʁ 2 VLAE 3 ࣮ݧɾ݁Ռ 6 / 25
ײతʹ... ͦͦɺRNN / AR ҙͷΛۙࣅͰ͖Δ 1 જࡏมʹใ͕΄ͱΜͲؚ·Εͳ͍ʢֶशॳظʣ 2 decoder σʔλΛ࠶ߏ͠Α͏ͱ͢Δ:
p(x|z) → pdecoder (x) 3 ࣄޙɾۙࣅࣄޙͱʹࣄલʹͳΔ p(z|x), q(z|x) → p(z) 7 / 25
গ͠ཧతʹ... VAE ≈ ූ߸Խ 1 σʔλͷຊ࣭ z Λූ߸Խ: p(z) 2
z ͷζϨΛූ߸Խ: p(x|z) ූ߸ͷ͞ʁ naive ʹ Cnaive (x) = Ex∼data,z∼q(z|x) [− log p(z) − log p(x|z)] Bits Back Coding ޮͷͨΊʹ encoder ͷ q(z|x) Λ༻͍Δ 8 / 25
Bits Back Coding q(z|x) ߴʑ H(q(z|x)) ϏοτͰใΛ͑ΒΕΔ ʢʣ ɿreceiver
q(z|x) ΛΈΕΔ߹ͷΈ Bits Back Coding ͷූ߸ Cnaive q(z|x) ͚ͩແବͰ L(x) = Eq(z|x) [log p(x|z) − log q(z|x)] ͳͷͰ CBitsBack (x) = Ex∼data [−L(x)] ≥ H(data) + Ex∼data [DKL (q(z|x)∥p(z|x))] 9 / 25
Bits Back Coding • ූ߸ͷ࠷খԽ = มԼքͷ࠷େԽ → z ͕ΘΕΔͷූ߸Խ͕ޮՌతͳ࣌
• ΑΓਖ਼֬ͳࣄޙʹΑΓมਪߴਫ਼ʹͳ Δ͕ɺݱ࣌Ͱଘࡏ͠ͳ͍ → DKL (≥ 0) ແࢹͰ͖ͳ͍ 10 / 25
Information Preference z ͕ແࢹ͞ΕΔͷ... p(x|z) ͕ pdata (x) Λz ͷใͳ͠ʹϞσϧԽͰ͖Δ߹
1 ࣄޙ pz|x) ͕ p(z) ʹͳΓɺ 2 ۙࣅࣄޙ q(z|x) p(z) ʹͳΔ ∵ KL ߲Λখ͘͢͞ΔͨΊ Information Preference • z ͳ͠ͰہॴతʹϞσϧԽͰ͖Δใہॴతʹ ූ߸Խ • ͦΕҎ֎ͷใ z Λͬͯ෮߸Խ જࡏมΛ hack ͢Δํ๏ɿ free bits, annealing the relative weight of DKL 11 / 25
1 ͳͥʁ 2 VLAE 3 ࣮ݧɾ݁Ռ 12 / 25
Ϟσϧͷ֓ཁ 1 ॊೈͳࣄલ 2 දݱྗͷ͋Δ decoder 13 / 25
ࣄલͷվળ • ٿ໘ΨεɾҰ༷͕ద͔ٙ • જࡏมͷ׆༻ʹෆՄܽ • → autoregressive flow 14
/ 25
Autoregressive Flow normalizing flows ʹ͍ͭͯ • ୯७ͳ͔ΒॊೈͳͷՄٯͳม • general normalizing
flow • volume preserving flow • Jacobian ͷѻ͍ʹҧ͍ AF ͷಛ IAF ͱಉ͡ܭࢉྔ͕ͩϞσϧ͕ΑΓਂ͍ 15 / 25
Inverse Autoregressive Flow zt = µt + σt ⊙ zt−1
log q(zT |x) = − D ∑ i=1 1 2 ϵ2 i + 1 2 log(2π) + T ∑ t=0 log σt,i ਤ 1: IAF ͷ֓ཁ 16 / 25
IAF posterior ॊೈͳࣄޙΛ֫ಘ͍ͯ͠Δʂ ਤ 2: IAF ͷࣄޙ 17 / 25
AF prior ≡ IAF posterior L(x; θ) = Ez∼q(z|x) [log
p(x|z) + log p(z) − log q(z|x)] = Ez∼q(z|x),ϵ=f−1(z) [ log p(x|f(ϵ)) + log u(ϵ) + log det dϵ dz − log q(z|x) ] = Ez∼q(z|x),ϵ=f−1(z) log p(x|f(ϵ)) + log u(ϵ) − ( log q(z|x) − log det dϵ dz ) IAF posterior 18 / 25
1 ͳͥʁ 2 VLAE 3 ࣮ݧɾ݁Ռ 19 / 25
࣮ݧ֓ཁ • త • જࡏม͕େҬతͳใΛ֫ಘ͍ͯ͠Δ͔ • AF prior ͕ IAF
posterior ΑΓ༏Ε͍ͯΔ͔ • AR decoder ʹΑΓີਪఆͷਫ਼্͕͕Δ͔ • ݕূϞσϧ: AF prior & PixelCNN decoder • σʔληοτ: 2 ͷ 28×28 ը૾ • MNIST, OMNIGLOT, Caltech - 101 Silhouettes • ΞʔΩςΫνϟɾજࡏมͷ࣍ݩ౷Ұ 20 / 25
Lossy Compression - MNIST ࠨɿೖྗɺӈɿग़ྗ • Ͳͷࣈ͔Θ͔Δ • ͨͩͷ࠶ߏͰͳ͍ ਤ
3: original & decompressed MNIST 21 / 25
Lossy Compression - OMNIGLOT ࠨɿೖྗɺӈɿग़ྗ • semantics ͕อଘ͞Ε ͍ͯͳ͍ •
λεΫɾσʔληοτ ͝ͱʹใΛಛఆ͢Δ ඞཁ ਤ 4: original & decompressed OMNIGLOT 22 / 25
જࡏม͔ΒͷαϯϓϦϯά • Սۭͷࣈ • େҬతͳಛ ਤ 5: VLAE ͔Βͷαϯϓϧ 23
/ 25
Density Estimation Unconditional Decoder γϯϓϧͳ PixelCNN 24 / 25
AF priorͷޮՌ • ີਪఆ͕վળ • AR ʹΑͬͯજࡏม ͷ࣋ͭใ͕૿Ճ ਤ 6:
AF prior ͷޮՌ 25 / 25