Upgrade to Pro
— share decks privately, control downloads, hide ads and more …
Speaker Deck
Sign up for free
Menu
Search
Features
All features
Private URLs
Password Protection
Custom URLS
Scheduled publishing
Remove Branding
Restrict embedding
Deck Collections
Notes
Features
All features
Private URLs
Password Protection
Custom URLS
Scheduled publishing
Remove Branding
Restrict embedding
Deck Collections
Notes
Explore
Featured decks
Featured speakers
Programming
Technology
Storyboards
Explore
Featured decks
Featured speakers
Programming
Technology
Storyboards
Pricing
Search
Sign in
Sign up for free
ACL読み会2020_Jointly Masked Sequence-to-Sequence ...
Search
maskcott
August 07, 2020
Research
30
0
Share
Embed
Copy iframe code
Copy JS code
Copy link
Start on current slide
ACL読み会2020_Jointly Masked Sequence-to-Sequence Model for Non-Autoregressive Neural Machine Translation
maskcott
August 07, 2020
More Decks by maskcott
See All by maskcott
論文紹介2022後期(EMNLP2022)_Towards Opening the Black Box of Neural Machine Translation: Source and Target Interpretations of the Transformer
maskcott
0
86
論文紹介2022後期(ACL2022)_DEEP: DEnoising Entity Pre-training for Neural Machine Translation
maskcott
0
46
PACLIC2022_Japanese Named Entity Recognition from Automatic Speech Recognition Using Pre-trained Models
maskcott
0
52
WAT2022_TMU NMT System with Automatic Post-Editing by Multi-Source Levenshtein Transformer for the Restricted Translation Task of WAT 2022
maskcott
0
60
論文紹介2022前期_Redistributing Low Frequency Words: Making the Most of Monolingual Data in Non-Autoregressive Translation
maskcott
0
71
論文紹介2021後期_Analyzing the Source and Target Contributions to Predictions in Neural Machine Translation
maskcott
0
90
WAT2021_Machine Translation with Pre-specified Target-side Words Using a Semi-autoregressive Model
maskcott
0
65
NAACL/EACL読み会2021_NEUROLOGIC DECDING: (Un)supervised Neural Text Generation with Predicate Logic Constraints
maskcott
0
52
論文紹介2021前期_Bilingual Dictionary Based Neural Machine Translation without Using Parallel Sentences
maskcott
0
60
Other Decks in Research
See All in Research
論文読み会 SNLP2026 Tau2-Bench: Evaluating Conversational Agents in a Dual-Control Environment
s_mizuki_nlp
0
240
Anthropic が提案する LLM の内部状態を自然言語で説明可能にした Natural Language Autoencoders / Natural Language Autoencoders Produce Unsupervised Explanations of LLM Activations
shunk031
0
210
VRID: View-Invariant Representation through Dual-Axis Transformation for Cross-iew Pose Estimation
satai
3
100
重要だけど測れていないもの:高齢者ケアの見えない課題
theoriatec2024
0
510
MM-OVSeg: Multimodal Optical–SAR Fusion for Open-Vocabulary Segmentation in Remote Sensing
satai
3
130
「AIとWhyを深堀る」をAIと深堀る
iflection
0
610
EIRによる不正端末のブロッキング 5G時代におけるデバイス識別と不正対策の進化
stellarcraft
0
130
Apache Gravitinoで実現する Icebergカタログ統合とアクセスの一元化
matsumooon
0
500
PHTalks Bengaluru - SSRF When All Else Fails
dk999
0
1.1k
敵対生成プロンプト同時探索による内省型プロンプト最適化
kinoue_smarthr
0
410
NLP colloquium: AI Safety Survey
kanekomasahiro
2
1.1k
Cross-Media Human-Information Interaction
signer
PRO
0
230
Featured
See All Featured
Refactoring Trust on Your Teams (GOTO; Chicago 2020)
rmw
35
3.8k
コードの90%をAIが書く世界で何が待っているのか / What awaits us in a world where 90% of the code is written by AI
rkaga
63
45k
Building Experiences: Design Systems, User Experience, and Full Site Editing
marktimemedia
0
600
The Impact of AI in SEO - AI Overviews June 2024 Edition
aleyda
6
1.2k
Future Trends and Review - Lecture 12 - Web Technologies (1019888BNR)
signer
PRO
0
3.7k
Building Better People: How to give real-time feedback that sticks.
wjessup
370
20k
brightonSEO & MeasureFest 2025 - Christian Goodrich - Winning strategies for Black Friday CRO & PPC
cargoodrich
3
820
Designing Powerful Visuals for Engaging Learning
tmiket
1
530
How to Build an AI Search Optimization Roadmap - Criteria and Steps to Take #SEOIRL
aleyda
1
2.2k
Product Roadmaps are Hard
iamctodd
55
13k
We Are The Robots
honzajavorek
0
350
Game over? The fight for quality and originality in the time of robots
wayneb77
1
270
Transcript
1
Abstract ・様々な自然言語処理のタスクで注目を集めているmasked languageモデルをseq2seqモデルに適用し た ”jointly masked seq2seq” モデルを提案しNATに適用した ・具体的にはトレーニング時にエンコーダーへの入力をマスキングし、デコーダーではn-gramのロス関数 で連続的にマスキングすることで学習するもの
・ WMT14 en-de/de-en で 27.69/32.24 のBLEUスコアを達成し、自己回帰モデルの5倍の速度を実現 2
Introduction ・NATモデルの精度が落ちる理由として次の二つが先行研究で主に挙げられている 1.ソース側の情報が適切にエンコードされていないこと 2.デコーダーがタスクをうまく処理できず、繰り返しや長文における性能が低下したりする →NATモデルのエンコーダーとデコーダーの機能を実験的に研究するとエンコーダーの方がデコーダーよ りも翻訳結果に影響を与えることが判明 ・BERTに倣ってエンコーダ―でマスキングを行うことでエンコーダーを徹底的に学習させる ・デコーダーの入力に対して連続的なマスキングを行う手法とn-gramロス関数の実装を提案 ・二つの方法を統合してjointly masked
seq2seqモデルを実装 3
Related Work ・Non-Autoregressive Machine Translation ターゲット文の文脈情報を捨てて全トークンを独立に出力することによって、ターゲット文の文長nに依存 しない、O(k)の計算量での翻訳が可能になった(kは定数) ・Masked Language Model
・BERT (Devlin et al., 2018)で提案され、トランスフォーマーのエンコーダー側を扱うモデル ・XLM (Lample and Conneau, 2019)では、ソース文とターゲット文をコンキャットしてエンコーダーの入 力とすることでクロスリンガルな情報を学習した ・MASS (Song et al., 2019)ではseq2seqにおける事前学習が提案されたが、モノリンガルなフレーム ワークだった →この論文ではATモデルのaccuracyとNATの推論速度を維持できるようなモデルに基づいた seq2seq のフレームワークでクロスリンガルな情報を扱う 4
Preliminary Study NATにおけるエンコーダーデコーダーモデルの構造を探るための実験的な研究 Nonautoregressive neural machine translation(Gu et al., 2017)で提案されたベーシックなNATモデル
を利用 データセット: IWSLT14 German to English エンコーダーとデコーダーの重要性を3つの観点から調べる 5
Preliminary Study 6
Methodology 問題設定 ・ソース文とターゲット文 ロス関数 エンコーダーマスキング 個の単語をランダムに選択して とする のうち80%を[mask], 10%をvocabからランダムに別の単語に置き換える 置き換えた後の文を とする
ロス関数 7
Methodology デコーダーマスキング ターゲット文 が与えられたときに、連続したトークンをマスキング エンコーダー同様にマスキングしたn-gramを , マスキング後の文を とする ロス関数 8
Methodology 連続的なn-gramベースのロス関数(Ma et al., 2018; Shao et al., 2018, 2019)も利用(デコーダー)
与えられるn-gram 9
Methodology 目的関数 デコード方法 エンコーダーの入力に特殊なトークンを加えて、そのトークンに相当する隠れベクトルから文長を予測する (Ghazvininejad et al., 2019) この文長に関するロスも上式に加えて計算 文長が決まったらターゲット文を[mask]で初期化してデコーダーに入力
出力のうち確率の低かった単語を選び、隣接する単語とともにマスクをしてデコーダーに入力を繰り返す (選ぶ単語の数は線形関数的に減衰させる) 事前に決めたイテレーション数回すか結果が変わらなくなったら終了 10
実験 データセット IWSLT14 German→English WMT16 English↔Romanian WMT14 English↔German Moses (Koehn
et al.,2007)でトークナイズ, Byte-Pair Encoding (BPE) (Sennrich et al., 2015)をかけて ソース文、ターゲット文で32kのvocabularyになった モデル Transformer (Vaswani et al., 2017) (IWSLTにはsmall, その他にはbase) NATモデル O(1)で推論できるものが5種類、O(k)で推論できるのを2種類先行研究から用意 sequence-level knowledge distillation(Kim and Rush, 2016)を適用 11
Result 12 O(1)のNAT O(k)のNAT
Analysis 13 encoder decoder
Ablation 14
Conclusion NATモデルの機能を実験的に調べ、エンコーダ―の学習の重要性を発見した エンコーダーの入力にマスキングを施しロス関数に基づく予測を提案することでエンコー ダ―の学習を向上させた デコーダー側ではn-gram単位のマスキングとn-gramロス関数を提案し、連続して出力 してしまう問題を和らげた 比較対象のベースラインにした全てのNATモデルよりも提案手法は優れたスコアを出 し、ATモデルの5倍以上の推論速度を達成した 15