Upgrade to Pro
— share decks privately, control downloads, hide ads and more …
Speaker Deck
Sign up for free
Menu
Search
Features
All features
Private URLs
Password Protection
Custom URLS
Scheduled publishing
Remove Branding
Restrict embedding
Deck Collections
Notes
Features
All features
Private URLs
Password Protection
Custom URLS
Scheduled publishing
Remove Branding
Restrict embedding
Deck Collections
Notes
Explore
Featured decks
Featured speakers
Programming
Technology
Storyboards
Explore
Featured decks
Featured speakers
Programming
Technology
Storyboards
Pricing
Search
Sign in
Sign up for free
NeuralIPS
Search
Wonseok Jung
December 11, 2018
440
0
Share
Embed
Copy iframe code
Copy JS code
Copy link
Start on current slide
NeuralIPS
Wonseok Jung
December 11, 2018
More Decks by Wonseok Jung
See All by Wonseok Jung
Ai for business -self car driving
wonseokjung
0
230
reinforcement_learning_.pdf
wonseokjung
2
1.8k
원석이의 모두연에서 강화학습 보석되기
wonseokjung
0
460
Introduction Deep Reinforcement Learning
wonseokjung
0
180
Deep reinforcemenet learning -2
wonseokjung
0
220
Deep Reinforcement Learning - Introduction
wonseokjung
1
680
How to become a datascientist ?
wonseokjung
2
2.4k
Review of Taylor series
wonseokjung
1
140
꿈꾸는 Agent
wonseokjung
2
170
Featured
See All Featured
Writing Fast Ruby
sferik
630
63k
RailsConf & Balkan Ruby 2019: The Past, Present, and Future of Rails at GitHub
eileencodes
141
35k
Principles of Awesome APIs and How to Build Them.
keavy
128
18k
Documentation Writing (for coders)
carmenintech
77
5.5k
WCS-LA-2024
lcolladotor
0
830
Avoiding the “Bad Training, Faster” Trap in the Age of AI
tmiket
0
230
Designing for Timeless Needs
cassininazir
1
470
Mobile First: as difficult as doing things right
swwweet
225
10k
Deep Space Network (abreviated)
tonyrice
0
300
SEOcharity - Dark patterns in SEO and UX: How to avoid them and build a more ethical web
sarafernandez
0
270
The Mindset for Success: Future Career Progression
greggifford
PRO
0
490
The Limits of Empathy - UXLibs8
cassininazir
1
660
Transcript
3-JOWJUFEUBML/FVSBM*14 8POTFPL+VOH
None
None
None
43࠙ীࢲ gradient descent methodۄҊ ೞחؘ, gradient ascent. rewardܳ maximize ೞח
policyܳ ӝ ਤ೧ࢲח gradient aascent .
43࠙ীࢲ gradient descent methodۄҊ ೞחؘ, gradient ascent. rewardܳ maximize ೞח
policyܳ ӝ ਤ೧ࢲח gradient aascent ܳ ೧ঠೠ.
None
None
None
None
None