Upgrade to Pro
— share decks privately, control downloads, hide ads and more …
Speaker Deck
Features
Speaker Deck
PRO
Sign in
Sign up for free
Search
Search
ICLR2017読み会@DeNA/iclr2017atDeNA_VLAE
Search
Masaki Kozuki
June 17, 2017
Research
29k
2
Share
Embed
Copy iframe code
Copy JS code
Copy link
Start on current slide
ICLR2017読み会@DeNA/iclr2017atDeNA_VLAE
Masaki Kozuki
June 17, 2017
More Decks by Masaki Kozuki
See All by Masaki Kozuki
Sanity Checks for Saliency Maps explained in Japanese language
crcrpar
0
2.7k
Deep Learning for clothes and changing pose
crcrpar
0
930
夏のトップカンファレンス論文読み会 / InnovationMeetup20170918csn_cvpr2k17
crcrpar
3
1.5k
iclr読み会 / iclrjp2017vlae
crcrpar
3
1.1k
Other Decks in Research
See All in Research
20260624 NLP colloquium: 単一のhubテキストがCLIPを壊す:hubnessによる埋め込みの脆弱性特定
de9uch1
2
250
GLIM とMegaParticles:正規分布近似の限界とタイトカップリング&パーティクルフィルタの進展 / GLIM and MegaParticles : Progress of the distribution representation in SLAM
koide3
0
810
JPA2026_NetworkTutorial_JunKashihara
junkashihara
0
130
【中間報告】国会議員の立法・政策実務を支える環境を巡る現状と課題
polipoli
0
570
CVPR2026論文紹介_VLMにとって良いvision encoderとは何か?Rethinking Model Selection in VLM Through the Lens of Gromov-Wasserstein Distance
kobayashi31
1
210
シングルチャネルマルチトーカー音声認識の進展
ryomasumura
0
260
LINEヤフー データサイエンス Meetup「三井物産コモディティ予測チャレンジ」の舞台裏-AlpacaTechパート
gamella
1
650
HAKARI-Bench - 実運用視点での情報検索モデル評価ベンチマーク
hotchpotch
1
730
VLMの推論を高速化する視覚トークン削減の仕組み
tattaka
2
300
大規模言語モデルは誰を覚えているか / Who Do Large Language Models Memorize?
upura
0
160
Karkada さんの論文 × 2 の紹介: (1) Closed-Form Training Dynamics Reveal Learned Features and Linear Structure in Word2Vec-like Models, (2) Symmetry in language statistics shapes the geometry of model representations
eumesy
PRO
1
700
論文紹介:Doc-to-LoRA: Learning to Instantly Internalize Contexts
yukako_nakano
0
150
Featured
See All Featured
Keith and Marios Guide to Fast Websites
keithpitt
413
23k
Let's Do A Bunch of Simple Stuff to Make Websites Faster
chriscoyier
508
140k
Music & Morning Musume
bryan
47
7.4k
Marketing to machines
jonoalderson
1
5.7k
The Success of Rails: Ensuring Growth for the Next 100 Years
eileencodes
47
8.3k
Designing Powerful Visuals for Engaging Learning
tmiket
1
530
Intergalactic Javascript Robots from Outer Space
tanoku
273
27k
Effective software design: The role of men in debugging patriarchy in IT @ Voxxed Days AMS
baasie
0
510
Imperfection Machines: The Place of Print at Facebook
scottboms
270
14k
The Mindset for Success: Future Career Progression
greggifford
PRO
0
490
What's in a price? How to price your products and services
michaelherold
247
13k
State of Search Keynote: SEO is Dead Long Live SEO
ryanjones
0
260
Transcript
Variational Lossy Autoencoder ICLR 2017 ಡΈձ @ DeNA @crcrpar 2017/6/17
1 / 25
จ • Variational Lossy Autoencoder • Xi Chen (UC Berkeley,
OpenAI), Diederik P. Kingma (OpenAI), Tim Salimans (OpenAI), et al. • දݱֶशͰજࡏมΛ׆༻͢Δ • Bits Back Coding Ͱ VAE ͷજࡏมʹ͍ͭͯͷߟ • જࡏมΛ lossy ʹ͢Δ • જࡏม z ͷ p(z), q(z|x) Λॊೈʹ • decoder ʹ PixelCNN 2 / 25
දهʹ͍ͭͯ • x ∈ Rd: σʔλ. x = ( x0
. . . xd )⊤ • x<i : x ͷ index ͕ i ະຬͷશཁૉ ( x0 . . . xi−1 )⊤ • z: જࡏม • pdata (x): σʔλΛੜ͢Δਅͷ • DKL (p∥q): p ͷ q ʹର͢Δ Kullback Leibler Divergence • θ: ϞσϧʢNNʣͷύϥϝʔλ • AR: PixelCNN ͳͲͷࣗݾճؼܕ NN • H, H: Τϯτϩϐʔ 3 / 25
VAE తؔ log p(X) = ∑ N i=1 log p(x(i))
࣮ࡍͷతؔ L(x; θ) = Eq(z|x) [log p(x|z) − DKL (q(z|x)∥p(z))] - ਖ਼نԽͨ͠ autoencoder ͱΈΕΔɻ VAE ͷ՝ɾऑ • දݱྗ͕ߴ͗͢Δ decoder જࡏมΛແࢹ • જࡏม͕ͭใΛཧͰ͖ͳ͍ 4 / 25
1 ͳͥʁ ײతʹ ཧʢBits Back Codingʣ 2 VLAE ֓ཁ Autoregressive
Flow decoder: PixelCNN 3 ࣮ݧɾ݁Ռ Lossy Comprssion Density Estimation 5 / 25
1 ͳͥʁ 2 VLAE 3 ࣮ݧɾ݁Ռ 6 / 25
ײతʹ... ͦͦɺRNN / AR ҙͷΛۙࣅͰ͖Δ 1 જࡏมʹใ͕΄ͱΜͲؚ·Εͳ͍ʢֶशॳظʣ 2 decoder σʔλΛ࠶ߏ͠Α͏ͱ͢Δ:
p(x|z) → pdecoder (x) 3 ࣄޙɾۙࣅࣄޙͱʹࣄલʹͳΔ p(z|x), q(z|x) → p(z) 7 / 25
গ͠ཧతʹ... VAE ≈ ූ߸Խ 1 σʔλͷຊ࣭ z Λූ߸Խ: p(z) 2
z ͷζϨΛූ߸Խ: p(x|z) ූ߸ͷ͞ʁ naive ʹ Cnaive (x) = Ex∼data,z∼q(z|x) [− log p(z) − log p(x|z)] Bits Back Coding ޮͷͨΊʹ encoder ͷ q(z|x) Λ༻͍Δ 8 / 25
Bits Back Coding q(z|x) ߴʑ H(q(z|x)) ϏοτͰใΛ͑ΒΕΔ ʢʣ ɿreceiver
q(z|x) ΛΈΕΔ߹ͷΈ Bits Back Coding ͷූ߸ Cnaive q(z|x) ͚ͩແବͰ L(x) = Eq(z|x) [log p(x|z) − log q(z|x)] ͳͷͰ CBitsBack (x) = Ex∼data [−L(x)] ≥ H(data) + Ex∼data [DKL (q(z|x)∥p(z|x))] 9 / 25
Bits Back Coding • ූ߸ͷ࠷খԽ = มԼքͷ࠷େԽ → z ͕ΘΕΔͷූ߸Խ͕ޮՌతͳ࣌
• ΑΓਖ਼֬ͳࣄޙʹΑΓมਪߴਫ਼ʹͳ Δ͕ɺݱ࣌Ͱଘࡏ͠ͳ͍ → DKL (≥ 0) ແࢹͰ͖ͳ͍ 10 / 25
Information Preference z ͕ແࢹ͞ΕΔͷ... p(x|z) ͕ pdata (x) Λz ͷใͳ͠ʹϞσϧԽͰ͖Δ߹
1 ࣄޙ pz|x) ͕ p(z) ʹͳΓɺ 2 ۙࣅࣄޙ q(z|x) p(z) ʹͳΔ ∵ KL ߲Λখ͘͢͞ΔͨΊ Information Preference • z ͳ͠ͰہॴతʹϞσϧԽͰ͖Δใہॴతʹ ූ߸Խ • ͦΕҎ֎ͷใ z Λͬͯ෮߸Խ જࡏมΛ hack ͢Δํ๏ɿ free bits, annealing the relative weight of DKL 11 / 25
1 ͳͥʁ 2 VLAE 3 ࣮ݧɾ݁Ռ 12 / 25
Ϟσϧͷ֓ཁ 1 ॊೈͳࣄલ 2 දݱྗͷ͋Δ decoder 13 / 25
ࣄલͷվળ • ٿ໘ΨεɾҰ༷͕ద͔ٙ • જࡏมͷ׆༻ʹෆՄܽ • → autoregressive flow 14
/ 25
Autoregressive Flow normalizing flows ʹ͍ͭͯ • ୯७ͳ͔ΒॊೈͳͷՄٯͳม • general normalizing
flow • volume preserving flow • Jacobian ͷѻ͍ʹҧ͍ AF ͷಛ IAF ͱಉ͡ܭࢉྔ͕ͩϞσϧ͕ΑΓਂ͍ 15 / 25
Inverse Autoregressive Flow zt = µt + σt ⊙ zt−1
log q(zT |x) = − D ∑ i=1 1 2 ϵ2 i + 1 2 log(2π) + T ∑ t=0 log σt,i ਤ 1: IAF ͷ֓ཁ 16 / 25
IAF posterior ॊೈͳࣄޙΛ֫ಘ͍ͯ͠Δʂ ਤ 2: IAF ͷࣄޙ 17 / 25
AF prior ≡ IAF posterior L(x; θ) = Ez∼q(z|x) [log
p(x|z) + log p(z) − log q(z|x)] = Ez∼q(z|x),ϵ=f−1(z) [ log p(x|f(ϵ)) + log u(ϵ) + log det dϵ dz − log q(z|x) ] = Ez∼q(z|x),ϵ=f−1(z) log p(x|f(ϵ)) + log u(ϵ) − ( log q(z|x) − log det dϵ dz ) IAF posterior 18 / 25
1 ͳͥʁ 2 VLAE 3 ࣮ݧɾ݁Ռ 19 / 25
࣮ݧ֓ཁ • త • જࡏม͕େҬతͳใΛ֫ಘ͍ͯ͠Δ͔ • AF prior ͕ IAF
posterior ΑΓ༏Ε͍ͯΔ͔ • AR decoder ʹΑΓີਪఆͷਫ਼্͕͕Δ͔ • ݕূϞσϧ: AF prior & PixelCNN decoder • σʔληοτ: 2 ͷ 28×28 ը૾ • MNIST, OMNIGLOT, Caltech - 101 Silhouettes • ΞʔΩςΫνϟɾજࡏมͷ࣍ݩ౷Ұ 20 / 25
Lossy Compression - MNIST ࠨɿೖྗɺӈɿग़ྗ • Ͳͷࣈ͔Θ͔Δ • ͨͩͷ࠶ߏͰͳ͍ ਤ
3: original & decompressed MNIST 21 / 25
Lossy Compression - OMNIGLOT ࠨɿೖྗɺӈɿग़ྗ • semantics ͕อଘ͞Ε ͍ͯͳ͍ •
λεΫɾσʔληοτ ͝ͱʹใΛಛఆ͢Δ ඞཁ ਤ 4: original & decompressed OMNIGLOT 22 / 25
જࡏม͔ΒͷαϯϓϦϯά • Սۭͷࣈ • େҬతͳಛ ਤ 5: VLAE ͔Βͷαϯϓϧ 23
/ 25
Density Estimation Unconditional Decoder γϯϓϧͳ PixelCNN 24 / 25
AF priorͷޮՌ • ີਪఆ͕վળ • AR ʹΑͬͯજࡏม ͷ࣋ͭใ͕૿Ճ ਤ 6:
AF prior ͷޮՌ 25 / 25