Upgrade to Pro
— share decks privately, control downloads, hide ads and more …
Speaker Deck
Features
Speaker Deck
PRO
Sign in
Sign up for free
Search
Search
Multi-Reference Training with Pseudo-References...
Search
ryoma yoshimura
January 23, 2019
Research
250
0
Share
Embed
Copy iframe code
Copy JS code
Copy link
Start on current slide
Multi-Reference Training with Pseudo-References for Neural Translation and Text Generation
研究室のEMNLP読み会の発表資料です。
ryoma yoshimura
January 23, 2019
More Decks by ryoma yoshimura
See All by ryoma yoshimura
TransQuest: Translation Quality Estimation with Cross-lingual Transformers
kokeman
0
290
Automatic Machine Translation Evaluation in Many Languages via Zero-Shot Paraphrasing
kokeman
0
65
BLEURT: Learning Robust Metrics for Text Generation
kokeman
0
270
Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks
kokeman
1
870
Courteously Yours: Inducing courteous behavior in Customer Care responses using Reinforced Pointer Generator Network
kokeman
0
180
Beyond BLEU: Training Neural Machine Translation with Semantic Similarity
kokeman
0
180
Reinforcement Learning Based Text Style Transfer without Parallel Training Corpus
kokeman
0
140
タスクとデータセット紹介 GLUE, SuperGLUE
kokeman
0
1.1k
Reliability and Learnability of Human Bandit Feedback for Sequence-to-Sequence Reinforcement Learning
kokeman
0
88
Other Decks in Research
See All in Research
セマンティック通信勉強会 6Gに向けたデバイス間効率的な通信の技術紹介・課題・今後展望
satai
3
280
EIRによる不正端末のブロッキング 5G時代におけるデバイス識別と不正対策の進化
stellarcraft
0
120
LINEヤフー データサイエンス Meetup「三井物産コモディティ予測チャレンジ」の舞台裏-AlpacaTechパート
gamella
1
630
SOTAのさらに先へ:厳しい推論制約下での高性能モデルのPost-Training
analokmaus
0
1.4k
NII S. Koyama's Lab Research Overview AY2026
skoyamalab
0
530
Anthropic が提案する LLM の内部状態を自然言語で説明可能にした Natural Language Autoencoders / Natural Language Autoencoders Produce Unsupervised Explanations of LLM Activations
shunk031
0
170
2025年度秋葉原ウォーカブルプロジェクト調査報告 「アキバらしいウォーカブル」とは何か
izumiyama_lab
1
180
さくらインターネット研究所テックトーク2026春、研究開発Gr.25年度成果26年度方針
kikuzo
0
180
VLMの推論を高速化する視覚トークン削減の仕組み
tattaka
2
250
AY 2026 Guide to Academic Writing Using Generative AI - Workshop
ks91
PRO
0
150
Scalable dynamic origin-destination demand estimation enhanced by high-resolution satellite imagery data
satai
3
410
2026年度 生成AI を活用した論文執筆ガイド/ワークショップ / 2026 Academic Year Guide to Writing Papers Using Generative AI - Workshop
ks91
PRO
0
210
Featured
See All Featured
Redefining SEO in the New Era of Traffic Generation
szymonslowik
1
390
Easily Structure & Communicate Ideas using Wireframe
afnizarnur
194
17k
Building Adaptive Systems
keathley
44
3.2k
Hiding What from Whom? A Critical Review of the History of Programming languages for Music
tomoyanonymous
3
1.1k
WCS-LA-2024
lcolladotor
0
800
Cheating the UX When There Is Nothing More to Optimize - PixelPioneers
stephaniewalter
287
14k
Navigating Team Friction
lara
192
16k
Dominate Local Search Results - an insider guide to GBP, reviews, and Local SEO
greggifford
PRO
0
300
The Illustrated Children's Guide to Kubernetes
chrisshort
51
53k
ピンチをチャンスに:未来をつくるプロダクトロードマップ #pmconf2020
aki_iinuma
128
56k
B2B Lead Gen: Tactics, Traps & Triumph
marketingsoph
0
220
Fight the Zombie Pattern Library - RWD Summit 2016
marcelosomers
234
17k
Transcript
Multi-Reference Training with Pseudo-References for Neural Translation and Text Generation
Renji Zheng, Mingbo Ma, Liang Huang EMNLP2018 研究室EMNLP読み会 紹介者 吉村
概要 • 複数のリファレンスでモデルを学習 ◦ テキスト生成の正解は1つではないので複数あったほうがいい ◦ 複数のリファレンスがあるデータセットを使用 • 複数のリファレンスから lattice
を作ってさらに多くの擬似リファ レンスを作成 ◦ 4~5個のリファレンスでは潜在的なリファレンスをカバーできない
Main Contributions • 機械翻訳と画像キャプションにおいてマルチリファレンスでの 学習法を3つ調査 • 複数の参照訳を lattice にするための新しいネットワークベー スの複数の系列アラインメントモデルを提案
• 擬似リファレンスでを用いた学習でMTでBLEUが+1.5、画像 キャプションでBLEUが+3.1、CIDErで+11.7
複数のリファレンスでの学習法 • 学習データを変えるだけでモデルは変更しなくていい • 複数のリファレンスがあるデータセットをシングルリファレンス のデータセットに変換 • 作り方はSample One、Uniform、Shuffleの3つ
複数のリファレンスでの学習法 • Sample One ◦ 各エポックでランダムに1つリファレンスを決める • Uniform ◦ 複数の各リファレンスに同じ入力をつける
• Shuffle ◦ Uniformで各エポックごとにシャッフルする x i : source y i : reference D : multiple reference dataset D’ : single reference dataset ※ D’ は順序集合
擬似リファレンスの作り方 • 複数のリファレンスから lattice を構築してそれをたどることで 擬似リファレンスを生成 ◦ 似た単語をマージする ◦ 元のリファレンスとBLEUを測って高いものを採用
• Hard alignと Soft align がある
Hard word Alignment • ペアワイズで同じ表層の単語をマージしていく • 以下の3文を考える
Hard word Alignment • Indonesia, its, opposition, foreign をマージ
Hard word Alignment • Indonesia, opposition, to, foreign をマージ •
(c)をたどることで 33個の擬似リファレンスができる
Hard Word Alignment の問題点 • 類義語を考慮できない ◦ 例での reiterated, repeats,
reiterates • 同一の単語は他の文では異なる意味をもつ可能性がある ◦ toなど(不定詞、前置詞)
Soft Word Alignment • 文y i と文y j に対して semantic
substitution matrix を作る • 各セルM u,v の値は単語y i,u と単語y j,v の類似度スコア • bidirectional LMの隠れベクトルのcos類似度 • Mを使ってアラインメントする ◦ M 0,0 からM |yi|,|yj| までの最適パスを動的計画法で求める
単語アラインメント 状態遷移関数 global penalty p: M u,v ≦ p では
align しない
Soft Word Alignment の結果
実験(MT) • NIST(2002-2005, 2006, 2008) zh-en ◦ single ref 1Mペア
(pre-train) 4 ref 5974ペア (train, valid, test) • global penalty 0.9 ◦ 100文集まるまで global penalty を減らしていく BLEUは上位50件のみ • bi-LMはpre-training dataとtraining dataで学習, word enmmbeding は Glove • encoderとdecoderは2層のbi-LSTMでBPEを使用 • pre-train: batch size 64, beam size 15, dropout 0.3 • multi-reference-train: batch size 100, 200, 400のベスト
Analysis of generated references • リファレンスの文長が長いほど、生成されるリファレンスの数が増える
結果
結果 各エポックで使うリファレンスの分散が高いため、 sample one はリファレンス数が10を越 えると急激に悪くなる
実験(Image Captioning) • MSCOCO • Resnet を LSTM に繋げる •
batch size: 50, 250, 500, 1000 での最適なサイズ • beam size: 5 • global penalty: 0.6
Analysis of generated references • MTと比べてオリジナルのリファレンスが短いので質が低く、数も少ない
MTと違ってShuffleが良くなってる ⇨ 機械翻訳の参照よりも多様であるから Uniform だと1つのバッチ内でリファレンスの 分散が大きくなるとモデルに悪影響
Case Study BLEUが100だが オリジナルリファレンスと は異なる文 BLEUが0だが画像を説明 できている
Conclusion • マルチリファレンスでの学習方法を調査 • 既存のマルチリファレンスから擬似リファレンスを生成する手法を提案 • MTと画像キャプションの両タスクでベースラインを上回る