Upgrade to Pro
— share decks privately, control downloads, hide ads and more …
Speaker Deck
Sign up for free
Menu
Search
Features
All features
Private URLs
Password Protection
Custom URLS
Scheduled publishing
Remove Branding
Restrict embedding
Deck Collections
Notes
Features
All features
Private URLs
Password Protection
Custom URLS
Scheduled publishing
Remove Branding
Restrict embedding
Deck Collections
Notes
Explore
Featured decks
Featured speakers
Programming
Technology
Storyboards
Explore
Featured decks
Featured speakers
Programming
Technology
Storyboards
Pricing
Search
Sign in
Sign up for free
B3勉強会 第2回 N-gramの紹介
Search
phong3112
February 29, 2016
140
0
Share
Embed
Copy iframe code
Copy JS code
Copy link
Start on current slide
B3勉強会 第2回 N-gramの紹介
phong3112
February 29, 2016
More Decks by phong3112
See All by phong3112
A Pointwise Approach for Vietnamese Diacritics Restoration
phong3112
0
130
文献紹介 2016-06-24:Building a Large Syntactically-Annotated Corpus of Vietnamese
phong3112
0
160
Smoothing: Add-1
phong3112
0
190
B3勉強会 第9回 言語モデルの評価
phong3112
0
120
Featured
See All Featured
Measuring Dark Social's Impact On Conversion and Attribution
stephenakadiri
2
280
Hiding What from Whom? A Critical Review of the History of Programming languages for Music
tomoyanonymous
3
1.2k
JAMstack: Web Apps at Ludicrous Speed - All Things Open 2022
reverentgeek
1
600
個人開発の失敗を避けるイケてる考え方 / tips for indie hackers
panda_program
123
22k
End of SEO as We Know It (SMX Advanced Version)
ipullrank
3
4.4k
Optimizing for Happiness
mojombo
378
71k
SEOcharity - Dark patterns in SEO and UX: How to avoid them and build a more ethical web
sarafernandez
0
280
How to audit for AI Accessibility on your Front & Back End
davetheseo
0
540
Building Adaptive Systems
keathley
44
3.2k
We Analyzed 250 Million AI Search Results: Here's What I Found
joshbly
1
1.9k
How to build a perfect <img>
jonoalderson
1
6k
Public Speaking Without Barfing On Your Shoes - THAT 2023
reverentgeek
1
560
Transcript
B3勉強会 第2回 2016年1月14日 N-gramの紹介 自然言語処理研究室 B3 LY NAM PHONG
はじめに • 参考文献 1.自然言語処理の基礎 奥村 学 著 2.https://class.coursera.org/nlp/lecture/14 • 内容
• 1.言語モデル • 2.N-gramモデル
言語モデル • 言語モデルとは、テキストにおける確率ということである。 • 例えば、P(hôm nay trời đẹp) > P(trời
đẹp hôm nay) (Pは確率と考えられる) • 言語モデルは自然言語処理の中にいろいろな地域を応用している、機械翻訳、スペルチェック とか。 • 機械翻訳 ➢ P(high winds tonight) > P(large winds tonight) • スペルチェック ➢ The office is about 15 minuets from my house. ➢ P(about 15 minutes) > P(about 15 minuets)
言語モデル • 言語モデルの目的は文字の確率を計算することである。 • P(W) = P(w 1 ,w 2
,w 3 ,...,w k ) • 関係のタスク:次の言葉の確率を計算する。 • P(w k |w 1 ,w 2 ,w 3 ,w 4 ,...,w k-1 ) • Bayes法則: P(w 1 w 2 w 3 ...w k )=P(w 1 )*P(w 2 |w 1 )*P(w 3 |w 1 w 2 )*...*P(w k |w 1 w 2 ..w k-1 ) • 例:確率P(“Today is Monday”)を計算する。 • Bayes法則に基づく、下の式に示す。 • P(“Today is Monday”)=P(Today)*P(is|Today)*P(Monday|Today is) • => 普通のは、nの値がすごく大きいから、計算できない!
N-gramモデル • Markov仮定を用いて、確率は近似値を計算できる。 • P(w k |w 1 ,w 2
,...w k-1 )≒P(w k |w k-n ,...,w k-1 ) =>文字kの確率はn文字前から得られる。これはN-gramモデルと言われている。 • 例:1-gram (unigram): P(w 1 w 2 ...w k )≒P(w 1 )*P(w 2 )*...*P(w k ) • 2-gram (bigram): P(w 1 w 2 ...w k )≒P(w 1 )*P(w 2 |w 1 )*...*P(w k |w k-1 ) • 3-gram (trigram): P(w 1 w 2 ...w k )≒P(w 1 )*P(w 2 |w 1 )*P(w 3 |w 1 ,w 2 )*...*P(w k |w k-2 ,w k-1 ) • N-gramの問題点: – 実は、言語が長距離の依存関係であるので、違う意味の場合もある。 • 例:The computer which I had just put into the machine room on 5th floor crashed.
計算例 • 1-gramで、下の場合はどちらの確率値が一番高い? • P(I like ice cream) • P(the
the the the) • P(I go to class daily) • P( I daily go to class)
計算例 • 下のような5文からなる英語の品詞タグ付コーパスを考え、P(N|Det)とP(V|Det)を計算しな さい。 A/Det cat/N sat/V on/P the/Det mat/N.
A/Det girl/N read/V a/Det book/N. Boys/N play/V baseball/N. A/Det train/N runs/V. A/Det dog/N chases/V a/Det cat/N.
計算例 P(N|Det) = C(Det,N)/C(Det) = 7/7 = 1 P(V|Det) =
C(Det,V)/C(Det) = 0/7 = 0 (Cは頻度を表すことにする)