Upgrade to Pro
— share decks privately, control downloads, hide ads and more …
Speaker Deck
Sign up for free
Menu
Search
Features
All features
Private URLs
Password Protection
Custom URLS
Scheduled publishing
Remove Branding
Restrict embedding
Deck Collections
Notes
Features
All features
Private URLs
Password Protection
Custom URLS
Scheduled publishing
Remove Branding
Restrict embedding
Deck Collections
Notes
Explore
Featured decks
Featured speakers
Programming
Technology
Storyboards
Explore
Featured decks
Featured speakers
Programming
Technology
Storyboards
Pricing
Search
Sign in
Sign up for free
文献紹介_20171110_QRNN _ Quasi-Recurrent Neural Net...
Search
Sponsored
·
Your Podcast. Everywhere. Effortlessly.
Share. Educate. Inspire. Entertain. You do you. We'll handle the rest.
→
hrsma2i
November 10, 2017
Research
62
0
Share
Embed
Copy iframe code
Copy JS code
Copy link
Start on current slide
文献紹介_20171110_QRNN _ Quasi-Recurrent Neural Networks
文献紹介
hrsma2i
November 10, 2017
More Decks by hrsma2i
See All by hrsma2i
文献紹介_20181123_SeqGAN_ Sequence Generative Adversarial Nets with Policy Gradient
hrsma2i
0
96
文献紹介_20180622_MUNIT _ Multimodal Unsupervised Image-to-Image Translation
hrsma2i
0
110
文献紹介_20180518_Pixel-Level Domain Transfer
hrsma2i
0
71
文献紹介_20180420_CSN _ Learning Type-Aware Embeddings for Fashion Compatibility
hrsma2i
0
190
Other Decks in Research
See All in Research
マーケットストリート 社会実験2024 in 秋葉原ジャンク通り 調査報告書
izumiyama_lab
1
140
2025年度秋葉原ウォーカブルプロジェクト調査報告 「アキバらしいウォーカブル」とは何か
izumiyama_lab
1
210
SoftMatcha 2: 1兆語規模コーパスの超高速かつ柔らかい検索
e869120_sub
7
3.8k
【ローカルAIに向き合う展示会vol.2】液体時間定数型モジュールを用いた オリジナルの双方向エンコーダーモデルNexteraBERT 推論速度向上検討並びにダウンストリーム評価
rikkabotan7
0
190
Apache Gravitinoで実現する Icebergカタログ統合とアクセスの一元化
matsumooon
0
500
敵対生成プロンプト同時探索による内省型プロンプト最適化
kinoue_smarthr
0
410
横浜市長(山中氏)の言動にかかる第三者による調査報告書
y150saya
0
150
SAM3を用いたコマ・吹き出しの 領域検出と分割構造からの読み順推定
kzmssk
0
120
GLIM とMegaParticles:正規分布近似の限界とタイトカップリング&パーティクルフィルタの進展 / GLIM and MegaParticles : Progress of the distribution representation in SLAM
koide3
0
820
Karkada さんの論文 × 2 の紹介: (1) Closed-Form Training Dynamics Reveal Learned Features and Linear Structure in Word2Vec-like Models, (2) Symmetry in language statistics shapes the geometry of model representations
eumesy
PRO
1
710
HAKARI-Bench - 実運用視点での情報検索モデル評価ベンチマーク
hotchpotch
1
740
RS-Agent: Automating Remote Sensing Tasks through Intelligent Agent
satai
3
540
Featured
See All Featured
10 Git Anti Patterns You Should be Aware of
lemiorhan
PRO
659
62k
Heart Work Chapter 1 - Part 1
lfama
PRO
8
36k
Gemini Prompt Engineering: Practical Techniques for Tangible AI Outcomes
mfonobong
2
520
Jess Joyce - The Pitfalls of Following Frameworks
techseoconnect
PRO
1
410
Designing for humans not robots
tammielis
254
26k
Improving Core Web Vitals using Speculation Rules API
sergeychernyshev
21
1.6k
The SEO Collaboration Effect
kristinabergwall1
1
550
Context Engineering - Making Every Token Count
addyosmani
9
1.1k
Practical Tips for Bootstrapping Information Extraction Pipelines
honnibal
25
2.1k
Design of three-dimensional binary manipulators for pick-and-place task avoiding obstacles (IECON2024)
konakalab
0
580
The Mindset for Success: Future Career Progression
greggifford
PRO
0
490
Agile Actions for Facilitating Distributed Teams - ADO2019
mkilby
0
280
Transcript
James Bradbury∗, Stephen Merity∗ , Caiming Xiong & Richard Socher
Salesforce Research Palo Alto, California arXiv:1611.01576v2 [cs.NE] 21 Nov 2016 ICLR 2017 accepted 文献紹介 QRNN: QUASI-RECURRENT NEURAL NETWORKS
Abstract - QRNN = RNN processing like CNN - can
process sequential data in parallel - up to 16 times faster than LSTM in train/test - can make visual analysis of weights easy
Outline - Introduction - review of RNN/LSTM - Model (QRNN)
- Variants - Results - sentiment classification - language modeling - character-level machine translation - Conclusion - Reference
Introduction (review of RNN) - the standard model architecture for
deep learning approaches to sequence modeling tasks - sentence classification | word- and character-level language modeling | machine translation | question answering | image caption | time series forecasting
Introduction (review of RNN) - the network which has loop
arhictectures - RNN is very deep (causing gradient vanising) word2vec(“私”) 昨日の株価 (“の”:0.2, “は”:0.3, ...) 今日の株価の予測値
Introduction (review of RNN) - problem: not good at learning
very long sequences - document classification | character-level - why?: can’t deal with sequential data in parallel
Introduction (review of LSTM) - LSTM solves gradient vanising, using
memory cell - LSTM has 3 gates to control information flow
Introduction (review of LSTM) - forget gate to control long-term
information (in memory cell c)
Introduction (review of LSTM) - input gate to control current+short-time
information (in x and h(t-1))
Introduction (review of LSTM) - update memory cell, mixing the
current with the previos memory cell
- output gate to control current hidden-state information to the
next layer Introduction (review of LSTM)
- using a forget gate instead of an input gate
Introduction (variants of LSTM)
Model
Model (convolution component) “ズン”, “ドコ”, “きよし” ( 1, 0, 0,
) =“ズン” この例はone-hotだが word2vecというもっといい変換 を使う ズン, ズン, ズン, ドコ, きよし 時刻tの値を予測す るのに未来の時刻 t+1のデータを用い てはいけないので、 masked convolution
Model
bottle-neckになっていた前の層のhidden state h[t-1] を用いるのではなく、前の時刻 の入力x[t-1/2/...]を用いて並列処理を可能 にした。 Model (pooling component) LSTM
さらに、hidden state h に重みをかけずに渡していくので、各要素 の情報がごっちゃにならないので可視化しやすい。 ここは従来のLSTMと同じく逐次 計算するが、そんなに大して時 間かからない。
Model (pooling component) other type poolings (この論文では使われていない?) f-pooling ifo-pooling
Variants - Zoneout: Dropout for LSTM - skip-connection like DenseNet
- Attention for Encoder-Decoder
Experiments - Sentiment Classification (document binary-classification) - IMDb movie review
- 25,000 positive/negative reviews - Language Modeling (word-level prediction) - PTB: Penn Treebank - Character-level Machine Translatoin - IWST English-German spoken language translation task
Results (sentiment classification) - 小 batch_size, 長 seq_len に向いている(最大16倍早 かった。)
- training 時間は3倍早い
Results (sentiment classification) final layer’s hidden state
Results (language modeling)
Results (character-level machine translation) BLEU: upper is better http://unicorn.ike.tottori-u.ac.jp/2010/s072046/paper/graduation-thesis/node32.html
考察 - LSTMに精度で少し負けてしまった理由は、隠れ層の状態 h[t-1] ではなく、直前の 入力 x[t-1|t-2|,...]を使って近似したからと考えられる。 - 入力で、隠れ層の状態を近似する場合、使う、前の時刻の filter
size k を無限大ま で長くすれば一致する。(sentiment classificationのtaskではkを大きくしたら精度 上がった) - なので、filter-sizeを大きくすればいいが、そうすると、計算速度はどれほど落ちる のかが問題。
Conclusion - QRNN = RNN processing like CNN - can
process sequential data in parallel - up to 16x faster than LSTM in train/test - can make visual analysis of weights easy
Reference - LSTM - LSTMネットワークの概要 https://qiita.com/KojiOhki/items/89cd7b69a8a6239d67ca - わかるLSTM ~ 最近の動向と共に
https://qiita.com/KojiOhki/items/89cd7b69a8a6239d67ca - ニューラルネットワーク勉強会 http://isw3.naist.jp/~neubig/student/2015/seitaro-s/161025neuralnet_study_LSTM.pdf - conv の 3D図作成 - thinkercad https://www.tinkercad.com/ - QRNN - LSTMを超える期待の新星、QRNN https://qiita.com/icoxfog417/items/d77912e10a7c60ae680e - slideshare https://www.slideshare.net/DeepLearningJP2016/dlquasirecurrent-neural-networks?qid=a4ead77d-d8dd-458b-965c-5e53723d7757 &v=&b=&from_search=1 - pytorchでの公式実装 https://github.com/salesforce/pytorch-qrnn/blob/master/torchqrnn/qrnn.py