Upgrade to Pro
— share decks privately, control downloads, hide ads and more …
Speaker Deck
Features
Speaker Deck
PRO
Sign in
Sign up for free
Search
Search
自然言語処理研究室_第06週.pdf
Search
Sponsored
·
SiteGround - Reliable hosting with speed, security, and support you can count on.
→
takegue
February 14, 2014
Technology
190
0
Share
Embed
Copy iframe code
Copy JS code
Copy link
Start on current slide
自然言語処理研究室_第06週.pdf
Bayesian Sets
takegue
February 14, 2014
More Decks by takegue
See All by takegue
不自然言語の自然言語処理: コード補完を支える最新技術
takegue
1
920
つかわれるプラットフォーム 〜デザイン編〜@DPM#2
takegue
2
12k
カルチャーとエンジニアリングをつなぐ データプラットフォーム
takegue
4
6.6k
toC企業でのデータ活用 (PyData.Okinawa + PythonBeginners沖縄 合同勉強会 2019)
takegue
4
1.2k
Rettyにおけるデータ活用について
takegue
0
950
Sparse Overcomplete Word Vector Representations
takegue
0
260
Aligning Sentences from Standard Wikipedia to Simple Wikipedia
takegue
0
250
High-Order Low-Rank Tensors for Semantic Role Labeling
takegue
0
140
Dependency-based empty category detection via phrase structure trees
takegue
0
120
Other Decks in Technology
See All in Technology
強化学習「理論」入門
enakai00
3
3.6k
新しい SLO が良い感じにハマっている話
z63d
5
2.1k
メルカリのグローバルアプリで挑んだ AlloyDB 運用と課題解決の実践記
hatappi
0
240
Breaking the Seal: Static Deobfuscation of Compiled V8 JavaScript Bytecode Malware
hshrzd
0
660
修正PRを食べてレビュースキルが賢くなる:Claude Codeによる自己改善サイクル
yuyaumetsu
6
1.4k
Data Hubグループ 紹介資料
sansan33
PRO
0
3.1k
Agent 時代の Kaggle 展望 / kaggle-in-the-agentic-era
upura
1
670
Forza Horizon 6 のテレメトリ機能で 自動運転に使えそうな学習データを集める話
henjin0
0
150
Webアクセシビリティ入門 2026
recruitengineers
PRO
3
460
DatadogのBits Chatが開発組織にもたらしたもの / What Bits Chat Has Brought Us
sms_tech
0
120
取引先から届く 「セキュリティチェックシート」の読み解き方
kamadamakoto
0
140
AIは実装を速くする。では、私たちは何を今作るべきか?-立場を越えてリリースに向き合ったチーム開発の実践 / 20260801 Hiromi Nakaya and Naoki Takahashi
shift_evolve
PRO
3
450
Featured
See All Featured
Building a Modern Day E-commerce SEO Strategy
aleyda
45
9.2k
Mobile First: as difficult as doing things right
swwweet
225
10k
Joys of Absence: A Defence of Solitary Play
codingconduct
1
430
How to optimise 3,500 product descriptions for ecommerce in one day using ChatGPT
katarinadahlin
PRO
2
3.7k
Organizational Design Perspectives: An Ontology of Organizational Design Elements
kimpetersen
PRO
1
790
Easily Structure & Communicate Ideas using Wireframe
afnizarnur
194
17k
Prompt Engineering for Job Search
mfonobong
0
400
We Analyzed 250 Million AI Search Results: Here's What I Found
joshbly
1
1.7k
Money Talks: Using Revenue to Get Sh*t Done
nikkihalliwell
0
460
HU Berlin: Industrial-Strength Natural Language Processing with spaCy and Prodigy
inesmontani
PRO
0
630
Building a A Zero-Code AI SEO Workflow
portentint
PRO
0
660
sira's awesome portfolio website redesign presentation
elsirapls
0
320
Transcript
自然言語処理研究室 B3 Seminar 2013 年度 第6週 ~論文紹介「Bayesian Sets」~ 長岡技術科学大学 B3
竹野 峻輔
• Bayesian Sets Ghahramani, Z. & Heller, K. Bayesian sets.
NIPS 2, 22–23 (2005). Google Sets 入力:複数の単語 ⇒ 出力:複数の単語と関連度の高い単語 ex) banana, apple ⇒ grape http://enspire.cocolog-nifty.com/blog/2011/07/google-sets-542.html http://googlesystem.blogspot.jp/2012/11/google-sets-still-available.html http://google.about.com/od/blogs/ss/Google-Labs-Dropouts-And-Failures_9.htm ベイズ推論を使ったモデル化 ⇒ Bayesian Sets 必要な知識(Keywords): ベイジアンネットワーク 2014/2/14 自然言語処理研究室 2013年度 B3ゼミ 第3週 1.Introduction
2014/2/14 自然言語処理研究室 2013年度 B3コアタイム 第3週 2. Bayesian Sets 基本的な方針; どうやって入力セットDc
と同じクラスタを見つけるか? ―クラスタのヒントとなるのはDc だけ ⇒所属するクラスタを探すのは困難 ⇒要求に応じたクラスタ(clustering on demand)を探す. ≒入力セットと似たような単語を見つける ⇒ 単語に類似度の優劣がつけば良い ⇒ 入力セットの類似度ランキングを作成. 上位の単語を出力
2014/2/14 自然言語処理研究室 2013年度 B3コアタイム 第3週 2. Bayesian Sets 基本的なアルゴリズム: パラメータθのもとDのデータが生み出されていると仮定する;ベイズモデル
与えられたサブセットDc から一番 共起しやすい単語を出力する 降順に並び替え⇒出力 パラメータθの打消し ベイズの定理
2014/2/14 自然言語処理研究室 2013年度 B3コアタイム 第3週 2. Bayesian Sets 評価関数: ⇒対数とると自己相互情報量(PMI)
DC : 入力のサブセット, xはアイテム アイテムの尤度の影響を打ち消して,共起を計る ⇒p(・)は分布関数だが実際には1列の行列で表現
2014/2/14 自然言語処理研究室 2013年度 B3コアタイム 第3週 3. Sparse Binary Data where
• 多変数ベルヌーイモデルを仮定 • ハイパーパラメータ導入(α,β);階層ベイズ
2014/2/14 自然言語処理研究室 2013年度 B3コアタイム 第3週 3. Sparse Binary Data Γ関数;
階乗(n!)計算の一般化 xi,j の場合分けを行い,簡単化
2014/2/14 自然言語処理研究室 2013年度 B3コアタイム 第3週 4.Exponential Families ベータ分布は凄い!-
2014/2/14 自然言語処理研究室 2013年度 B3コアタイム 第3週 4.Exponential Families score関数を指数型分布での一般化 h :正規化関数,
ν :頻度分布 事前に計算できる部分と計算できない部分に分けれる ⇒ 計算速度の高速化
2014/2/14 自然言語処理研究室 2013年度 B3コアタイム 第3週 4.Exponential Families
2014/2/14 自然言語処理研究室 2013年度 B3コアタイム 第3週 5. Sparse Binary Data 結果:
Grolier Encyclopedia Data から作られたBayesian SetsとGoogle Setsの比較 素性抽出方法(2値の素性): (article, word)の組み合わせで素性抽出;ただし頻度1のデータは取り除く. α = cm, β = c(1-m) where m:平均ベクトル, c :定数(2) 応答時間:1.1秒程度 (MATLAB, Pentium4@2GHz