Upgrade to Pro
— share decks privately, control downloads, hide ads and more …
Speaker Deck
Sign up for free
Menu
Search
Features
All features
Private URLs
Password Protection
Custom URLS
Scheduled publishing
Remove Branding
Restrict embedding
Deck Collections
Notes
Features
All features
Private URLs
Password Protection
Custom URLS
Scheduled publishing
Remove Branding
Restrict embedding
Deck Collections
Notes
Explore
Featured decks
Featured speakers
Programming
Technology
Storyboards
Explore
Featured decks
Featured speakers
Programming
Technology
Storyboards
Pricing
Search
Sign in
Sign up for free
金研究室 勉強会 『Attention is all you need』
Search
Sponsored
·
Your Podcast. Everywhere. Effortlessly.
Share. Educate. Inspire. Entertain. You do you. We'll handle the rest.
→
winnie279
August 12, 2021
Science
170
0
Share
Embed
Copy iframe code
Copy JS code
Copy link
Start on current slide
金研究室 勉強会 『Attention is all you need』
Attention is all you need, Ashish et al., 2017, arXiv:1706.03762
winnie279
August 12, 2021
More Decks by winnie279
See All by winnie279
NowWay:訪⽇外国⼈旅⾏者向けの災害⽀援サービス
yjn279
0
41
「みえるーむ」(都知事杯Open Data Hackathon 2024 Final Stage)
yjn279
0
110
「みえるーむ」(都知事杯オープンデータ・ハッカソン 2024)
yjn279
0
99
5分で学ぶOpenAI APIハンズオン
yjn279
0
260
『確率思考の戦略論』
yjn279
0
170
Amazonまでのレコメンド入門
yjn279
1
210
もう一度理解するTransformer(後編)
yjn279
0
98
金研究室 勉強会 『もう一度理解する Transformer(前編)』
yjn279
0
150
金研究室 勉強会 『U-Netとそのバリエーションについて』
yjn279
0
1.1k
Other Decks in Science
See All in Science
[TMLR 2026, Featured Certification] Double Bounded α-Divergence Optimization for Density Estimation
gkazunii
1
120
データベース08: 実体関連モデルとは?
trycycle
PRO
0
1.6k
20260410_SystemsThinking
takusamar
1
160
Understanding CVP Waveforms: Interpretation and Clinical Implications in Anesthesiology
taka88
0
900
How a camera trap data standard enabled an ecosystem of interoperable tools
peterdesmet
0
130
O(log n)-Approximation Algorithms for Bipartiteness Ratio
tasusu
0
210
[NLP2026 参加報告会] AI for Science まとめ / NLP2026
lychee1223
0
2k
J-STAGE全文XML登載必須化について
xspa2012
0
1.5k
Visual Linear Algebra - Lecture at Shosen Grande
hiranabe
0
560
(SIGBIO84) Inverse MSMD法による化合物部分構造プロファイリングと結合親和性推定
keisukeyanagisawa
PRO
0
130
機械学習 - 授業概要
trycycle
PRO
0
650
2026 Introduction to University Math 01
kanaya
0
150
Featured
See All Featured
StorybookのUI Testing Handbookを読んだ
zakiyama
31
6.9k
30 Presentation Tips
portentint
PRO
1
410
A better future with KSS
kneath
240
18k
Claude Code どこまでも/ Claude Code Everywhere
nwiizo
67
58k
Music & Morning Musume
bryan
48
7.4k
Have SEOs Ruined the Internet? - User Awareness of SEO in 2025
akashhashmi
0
500
Making the Leap to Tech Lead
cromwellryan
135
10k
10 Git Anti Patterns You Should be Aware of
lemiorhan
PRO
659
62k
Leo the Paperboy
mayatellez
10
2.3k
ピンチをチャンスに:未来をつくるプロダクトロードマップ #pmconf2020
aki_iinuma
128
56k
The MySQL Ecosystem @ GitHub 2015
samlambert
251
13k
Java REST API Framework Comparison - PWX 2021
mraible
34
9.7k
Transcript
Attention Is All You Need Ashish et al., 2017, arXiv:1706.03762
金研 機械学習勉強会 2021/08/12 中村勇士
Transformerとは? • RNNの問題点 ◦ 長い入力が苦手 ◦ 勾配消失問題が起こりやすい ◦ 並列化が困難 →
GPUによる学習の効率化・大量のデータによる学習が困難 • Transformerによる解決 ◦ 再帰や畳み込みを使用しない ◦ 大規模なモデル・データを使用可能 ◦ 精度の大幅な向上
EQTransformerとの関係 • Transformerをそのまま使用していない ◦ attentionをレイヤーと使用 • 疑問 ◦ Transformerの強み: 再帰や畳み込みをしないこと
◦ LSTM・Convを使って良いのか?
モデル • エンコーダ・デコーダ • Attention • フィード・フォワード・ネットワーク(FFW) • 埋め込み •
ポジショナル・エンコーディング
モデル:エンコーダ・デコーダ
• エンコーダ(左) ◦ input ◦ N = 6 • デコーダ(右)
◦ output ◦ N = 6 モデル:エンコーダ・デコーダ input からの 出力
モデル:埋め込み / ポジショナル・エンコーディング • 埋め込み:単語のベクトル化 ◦ • ポジショナル・エンコーディング ◦ 構造のベクトル化
◦ 再帰や畳み込みの必要がなくなる ◦ モデルの学習が容易になる pos: 単語の順番, i: 次元, d model : 全体の次元数
モデル:Attention • 単語間の相関を表す ◦ どの単語がどの単語に 着目してるか • Q:query • K:key
• V:value • d k :dimention
Transformerの活用 • 自然言語処理(NLP) ◦ BERT ◦ GPT-3 ◦ DALL・E(テキストから画像生成) •
その他 ◦ 地震学:EQTransformer(地震動検出・フェーズピック) ◦ 生物学:AlphaFold2(タンパク質の構造予測) ◦ 音楽:Music Transformer(作曲)
おまけ • Transformer解説:GPT-3、BERT、T5の背後にあるモデルを理解する ◦ AINOW ◦ https://ainow.ai/2021/06/25/256107 • The Illustrated
Transformer ◦ Jay Alammar ◦ http://jalammar.github.io/illustrated-transformer • Embedding Projector ◦ http://projector.tensorflow.org/
モデル:フィード・フォワード・ネットワーク(FFW) • FFW ◦ 2つの線形変換 ◦ ReLU • 学習 ◦
英独:450万の文, 37,000のトークン ◦ 英仏: