Upgrade to Pro
— share decks privately, control downloads, hide ads and more …
Speaker Deck
Features
Speaker Deck
PRO
Sign in
Sign up for free
Search
Search
Metalearning shared Hierarchy
Search
Sponsored
·
Your Podcast. Everywhere. Effortlessly.
Share. Educate. Inspire. Entertain. You do you. We'll handle the rest.
→
Wonseok Jung
August 28, 2018
Science
0
51
Metalearning shared Hierarchy
Metalearning shared Hierarchy
논문 review
Wonseok Jung
August 28, 2018
Tweet
Share
More Decks by Wonseok Jung
See All by Wonseok Jung
Ai for business -self car driving
wonseokjung
0
210
reinforcement_learning_.pdf
wonseokjung
2
1.5k
원석이의 모두연에서 강화학습 보석되기
wonseokjung
0
430
NeuralIPS
wonseokjung
0
430
Introduction Deep Reinforcement Learning
wonseokjung
0
170
Deep reinforcemenet learning -2
wonseokjung
0
210
Deep Reinforcement Learning - Introduction
wonseokjung
1
660
How to become a datascientist ?
wonseokjung
2
2.3k
Review of Taylor series
wonseokjung
1
120
Other Decks in Science
See All in Science
【RSJ2025】PAMIQ Core: リアルタイム継続学習のための⾮同期推論・学習フレームワーク
gesonanko
0
710
コミュニティサイエンスの実践@日本認知科学会2025
hayataka88
0
140
Accelerating operator Sinkhorn iteration with overrelaxation
tasusu
0
240
Celebrate UTIG: Staff and Student Awards 2025
utig
0
1.3k
Algorithmic Aspects of Quiver Representations
tasusu
0
230
A Guide to Academic Writing Using Generative AI - A Workshop
ks91
PRO
0
240
20251212_LT忘年会_データサイエンス枠_新川.pdf
shinpsan
0
260
データマイニング - グラフデータと経路
trycycle
PRO
2
480
知能とはなにかーヒトとAIのあいだー
tagtag
PRO
0
180
MATSUO Makiko
genomethica
0
110
白金鉱業Meetup_Vol.20 効果検証ことはじめ / Introduction to Impact Evaluation
brainpadpr
2
1.7k
データマイニング - グラフ構造の諸指標
trycycle
PRO
0
280
Featured
See All Featured
GitHub's CSS Performance
jonrohan
1032
470k
How To Speak Unicorn (iThemes Webinar)
marktimemedia
1
410
Utilizing Notion as your number one productivity tool
mfonobong
4
270
The World Runs on Bad Software
bkeepers
PRO
72
12k
Mozcon NYC 2025: Stop Losing SEO Traffic
samtorres
0
190
How to make the Groovebox
asonas
2
2k
Side Projects
sachag
455
43k
The untapped power of vector embeddings
frankvandijk
2
1.6k
svc-hook: hooking system calls on ARM64 by binary rewriting
retrage
2
180
Deep Space Network (abreviated)
tonyrice
0
96
The Curse of the Amulet
leimatthew05
1
10k
How to build an LLM SEO readiness audit: a practical framework
nmsamuel
1
690
Transcript
.FUB-FBOJOHTIBSFE)JFSBSDIZ 8POTFPL+VOH 3FJOGPSDFNFOU-FBSOJOH
ਗࢳ 8POTFPL+VOH $JUZ6OJWFSTJUZPG/FX:PSL#BSVDI$PMMFHF %BUB4DJFODF.BKPS $POOFYJPO"*"*3FTFBSDIFS %FFQ-FBSOJOH$PMMFHF3FJOGPSDFNFOU-FBSOJOH3FTFBSDIFS .PEVMBCT$53--FBEFS 3FJOGPSDFNFOU-FBSOJOH 0CKFDU%FUFDUJPO
$IBUCPU (JUIVC IUUQTHJUIVCDPNXPOTFPLKVOH 'BDFCPPL IUUQTXXXGBDFCPPLDPNXTKVOH #MPH IUUQTXPOTFPLKVOHHJUIVCJP
ݾର 1. Introduction 2. Problem Statement 3. Algorithm 4. Experiments
META LEARNING SHARED HIERARCHIES
1.INTRODUCTION
1. UTILIZE PRIOR KNOWLEDGE META LEARNING SHARED HIERARCHIES 6UJMJ[FQSJPSLOPXMFEHF .BTUFSOFXUBTL
1.1 BUT REINFORCEMENT… META LEARNING SHARED HIERARCHIES How about Reinforcement
Learning?
1.2 SOLVE EACH TASK INDEPENDENTLY AND FROM SCRATCH SUPERMARIO WITH
R.L https://www.youtube.com/watch?v=IjvbhwuCaF0
1.3 ISSUES META LEARNING SHARED HIERARCHIES Sharing information Task1 Task2
Task3 θ1 θ2 θ3
1.4 MASTER POLICY META LEARNING SHARED HIERARCHIES Master Policy Sub1
Sub2 Sub3 θ1 θ2 θ3
1.5 MLSH META LEARNING SHARED HIERARCHIES Metalearning shared hierarchies
2.PROBLEM STATEMENT
2.1 NOTATION Time step Action Transition Function Reward Set of
states Set of actions Start state Discount factor t a P(s′, r ∣ s, a) r A S S0 γ Set of reward Policy Reward State R π r REINFORCEMENT LEARNING s
2.2 NOTATION META LEARNING SHARED HIERARCHIES EJTUSJCVUJPOPWFS.%1T "HFOUחQBSBNFUFSWFDUPSܳӝਵ۽VQEBUFೠ పझٜՙܻҕਬೞחۄఠ
пపझۄఠ BHFOUоഅపझ.ਸߓݴসؘೞחۄఠ PM πθ,ϕ(a∣s) ϕ θ
"DUJPO "HFOU &OWJSPONFOU 3FXBSE At Rt 4UBUF St Rt+1 St+1
REINFORCEMENT LEARNING 2.3 OBJECTIVE MDP
REINFORCEMENT LEARNING 2.4 NEW MDP &OWJSPONFOU 3FXBSE At Rt St
Rt+1 St+1 5BQUIFCBMM 1PTJUJWF3FXBSE New MDP
SUPERMARIO WITH R.L 2.5 NEW MDP-2 "DUJPO "HFOU &OWJSPONFOU 3FXBSE
At Rt 4UBUF St Rt+1 St+1 3FXBSE 1FOBMUZ Another New MDP
2.6 FIND SHARING PARAMETER META LEARNING SHARED HIERARCHIES maximizeϕ EM∼PM
, t = 0...T − 1[R]
2.7 STRUCTURE META LEARNING SHARED HIERARCHIES
3.ALGORITHM
3.1 MLSH ALGORITHM META LEARNING SHARED HIERARCHIES
3.2 MLSH ALGORITHM META LEARNING SHARED HIERARCHIES Two main components
3.3 MLSH ALGORITHM META LEARNING SHARED HIERARCHIES Joint update period
Warmup period
3.4 MLSH ALGORITHM META LEARNING SHARED HIERARCHIES Joint update period
Warmup period
3.5 MLSH ALGORITHM META LEARNING SHARED HIERARCHIES Joint update period
Warmup period θ θ, ϕ update
3.6 MLSH ALGORITHM-2 META LEARNING SHARED HIERARCHIES Joint update period
Warmup period θ θ, ϕ update
3.7 MLSH ALGORITHM-WARMUP META LEARNING SHARED HIERARCHIES update
3.8 MLSH ALGORITHM- JOINT UPDATE PERIOD META LEARNING SHARED HIERARCHIES
update
3.8 MLSH ALGORITHM META LEARNING SHARED HIERARCHIES update
4. EXPERIMENTS
4.1 2D MOVING BANDITS TASK META LEARNING SHARED HIERARCHIES
4.2 RESULT(2D BALL) META LEARNING SHARED HIERARCHIES
4.3 WALKING, CRAWLING META LEARNING SHARED HIERARCHIES
4.4 WALKING, CRAWLING META LEARNING SHARED HIERARCHIES