Word Embeddings 1. Introduction 2. Hiragana-Kanji conversion task 3. Our approach of WSD 4. The Effect of window size 5. The Effect of data size 1 #E15101
Word Embeddings 1. Introduction 2. Hiragana-Kanji conversion task 3. Our approach of WSD 4. The Effect of window size 5. The Effect of data size 2 #E15101
enabling computers to understand the sense of words. • Because of ambiguity, decline in accuracy occurs in many tasks such as machine translation. 4 cool awesome Ex.) That new bike is cool. not warm, but not cold Ex.) My ice-cream is cool.
et al. 2017] • The lack of groundbreaking improvements • The difficulty of integrating current WSD systems into downstream NLP applications 6 The system output result is not useful in downstream NLP applications
My ice-cream is cool. • System output : “senseid = 2” • Although this information is indeed useful, it is not easy to utilize this in the following procedures. 7
Hiragana- Kanji Conversion task [Yamamoto et al., 2016] • System input is sentence, and output is word. Not sense id. • Use the characteristics of Japanese language 8
size problem in Japanese WSD task. • New WSD method using word embeddings. • Changes in accuracy when changing the size of the training data. 10 In Japanese, there is no research about window size. Most of study use words that appear in two words to right and left.
size problem in Japanese WSD task. • New WSD method using word embeddings and PMI. • Changes in accuracy when changing the size of the training data. 11
size problem in Japanese WSD task. • New WSD method using word embeddings. • Changes in accuracy when changing the size of the training data. 12 There is no research about changes in accuracy when increase training data size due to cost problem.
Word Embeddings 1. Introduction 2. Hiragana-Kanji conversion task 3. Our approach of WSD 4. The Effect of window size 5. The Effect of data size 13 #E15101
and multiple corresponding Kanji words 18 I go to Tokyo and I buy a game in there. ࢲ౦ژʹߦͬͯήʔϜΛ͔͏ɻ ͔͏ LBV ങ͏ LBV UPCVZ TPNFUIJOH ࣂ͏ LBV UPIBWF QFUT Hiragana • phonographic • ambiguous Kanji • ideographic • clear sense
regarded as a WSD task. • We can get Kanji word as a WSD result and we can use it in the following procedures without any changes. 19 ͔͏ LBV ങ͏ LBV UPCVZ TPNFUIJOH ࣂ͏ LBV UPIBWF QFUT Hiragana • phonographic • ambiguous Kanji • ideographic • clear sense
The ease of making data sets We can make data sets from large amount of raw corpus automatically since convert Hiragana to Kanji is very easy. • The ease of analysis Since we can create data sets freely, we can create training data that matches the problem we want to investigate. Ex. ) Window size problem, Data size problem 20
The ease of making data sets We can make data sets from large amount of raw corpus automatically since convert Hiragana to Kanji is very easy. • The ease of analysis Since we can create data sets freely, we can create training data that matches the problem we want to investigate. Ex. ) Window size problem, Data size problem 21
The ease of making data sets We can make data sets from large amount of raw corpus automatically since convert Hiragana to Kanji is very easy. • The ease of analysis Since we can create data sets freely, we can create training data that matches the problem we want to investigate. Ex. ) Window size problem, Data size problem 22
Word Embeddings 1. Introduction 2. Hiragana-Kanji conversion task 3. Our approach to WSD 4. The Effect of window size 5. The Effect of data size 23 #E15101
AVE that employ word embeddings. [Sugawara et al., 2015] • The similarity in CWE is not reflected depending on the appearance position, because the word position is restricted. 26 CWE: AVE:
CWE+AVE+PMI is a method that joins vectors of the word with the maximum PMI as features. • PMI is calculated between the target Kanji and all the words in the sentence. 28
Word Embeddings 1. Introduction 2. Hiragana-Kanji conversion task 3. Our approach to WSD 4. The Effect of window size 5. The Effect of data size 30 #E15101
Corpus of Contemporary Written Japanese (BCCWJ) • 200 sentences per sense (Kanji) were sampled from BCCWJ as training data. • The corpus used for constructing word embedding was created from the data of Japanese Wikipedia. • LinearSVC implemented by scikit-learn was used as a classifier. 31
deriving the meaning of “playing” from the musical term “Adagio” • Adding AVE as a feature, CWE+AVE solved the problem caused by a restriction in word position and derived the correct Kanji. 33
the word “idea” is added as the feature. • The CWE+AVE method does not utilize “idea” which introduces “devote” in window size 10; nevertheless, this is solved by using PMI which adds “idea” ’s word embedding. 34
Word Embeddings 1. Introduction 2. Hiragana-Kanji conversion task 3. Our approach to WSD 4. The Effect of window size 5. The Effect of data size 35 #E15101
words that appear in the whole sentence. • We also show that the data required to obtain high accuracy increases exponentially. • Future work: developing the tool for WSD with higher accuracy. 38
• There are 28,467,950 Hiragana words. • BCCWJ contain 438,360 Hiragana words that are target of Hiragana-Kanji conversion task. • One hiragana word appears per about 13 sentences. 39
CWE+AVE+PMI is achieved about two percentage points than CWE. • Considering 5-7 or more window size is important for accurate Word Sense Disambiguation. 40