Upgrade to Pro
— share decks privately, control downloads, hide ads and more …
Speaker Deck
Speaker Deck
PRO
Sign in
Sign up for free
機械学習を⽤いた⽇経電⼦版Proのユーザ分析 / Data Analysis in Nikkei using Machine Learning
Shotaro Ishihara
January 22, 2019
Business
8
10k
機械学習を⽤いた⽇経電⼦版Proのユーザ分析 / Data Analysis in Nikkei using Machine Learning
Data Driven Developer Meetup #4 (#d3m) での発表資料
https://d3m.connpass.com/event/115217/
Shotaro Ishihara
January 22, 2019
Tweet
Share
More Decks by Shotaro Ishihara
See All by Shotaro Ishihara
Analysis and Estimation of News Article Reading Time with Multimodal Machine Learning
upura
0
79
データ分析の進め方とニュースメディアでのデータ活用事例 / data-analysis-in-kaggle-and-news-media
upura
0
470
国際会議参加報告 AACL-IJCNLP 2022 / AACL-IJCNLP 2022 Report
upura
0
260
spoana 2022年の活動報告と 来年以降の企画募集 / spoana-2022
upura
0
490
Semantic Shift Stability: Efficient Way to Detect Performance Degradation of Word Embeddings and Pre-trained Language Models
upura
0
820
実践:日本語文章生成 Transformers ライブラリで学ぶ実装の守破離 / Introduction of Japanese Text Generation with Transformers
upura
5
7.5k
Nikkei at SemEval-2022 Task 8: Exploring BERT-based Bi-Encoder Approach for Pairwise Multilingual News Article Similarity
upura
0
420
[Poster] Nikkei at SemEval-2022 Task 8: Exploring BERT-based Bi-Encoder Approach for Pairwise Multilingual News Article Similarity
upura
0
380
新聞記事のクリック率予測に向けたペアワイズ学習用データセットの構築手法の検討 / JSAI2022 Generating Pairwise Dataset for CTR Prediction
upura
0
85
Other Decks in Business
See All in Business
20230122_BTCONJP_C15
barusu
0
110
Epics - Buidl to Earn.
epicsdao
1
280
20230118 kazaneya TeckTalk3 Data Standards and Open Data Initiatives by the Digital Agency of Japan
haseryo
5
3.8k
レンティオ株式会社 採用候補者様向け会社紹介資料 / Company Profile
rentio
PRO
0
5.3k
X Mile人事制度丸わかりBook
xmile
PRO
0
1.2k
tetemarche recruite
tetemarche
1
7.4k
MNTSQ CompanyDeck
mntsq
0
17k
AKIBA.SaaS #3 - No.4
cmsakumashogo
0
160
RPALT Gundam vol.2 / LT4 : The result of enjoying First Gundam with an RPA brain
tsutomu_asari
0
370
Contrea Company Deck
contrea_0123
3
6.1k
Twitter Instagram キャンペーンツール「キャンつく」
pickles_staff
0
44k
やる人・やり続ける人・やり切る人の「大きな違い」がわかる資料
nyattx
PRO
1
1.1k
Featured
See All Featured
Three Pipe Problems
jasonvnalue
89
8.9k
Keith and Marios Guide to Fast Websites
keithpitt
407
21k
CoffeeScript is Beautiful & I Never Want to Write Plain JavaScript Again
sstephenson
152
13k
The Power of CSS Pseudo Elements
geoffreycrofte
52
4.3k
RailsConf & Balkan Ruby 2019: The Past, Present, and Future of Rails at GitHub
eileencodes
120
29k
Statistics for Hackers
jakevdp
784
210k
Building Your Own Lightsaber
phodgson
96
4.9k
Gamification - CAS2011
davidbonilla
75
4.1k
GraphQLとの向き合い方2022年版
quramy
20
9.8k
The MySQL Ecosystem @ GitHub 2015
samlambert
240
11k
Understanding Cognitive Biases in Performance Measurement
bluesmoon
2
390
Helping Users Find Their Own Way: Creating Modern Search Experiences
danielanewman
10
1.3k
Transcript
ػցֶशΛ༻͍ͨ ܦిࢠ൛1SPͷϢʔβੳ ຊܦࡁ৽ฉࣾ ੴݪↅଠ %BUB%SJWFO%FWFMPQFS.FFUVQ +BOOE
ٕज़ॻయͰࣥචɾެ։ ٕज़ॻయ̑Ͱ൦ͨ͠ܦిࢠ൛ͷٕज़ॻΛ࠶ൢ͠·͢ɻ IUUQTOPUFNV
[email protected]
OODCBC • ୲ͨ͠ୈষʮػցֶशΛ༻͍ͨܦిࢠ൛1SP ͷϢʔβੳʯશͯແঈެ։த
ຊͷ • ࣗݾհ • σʔλಓͱʮܦిࢠ൛1SPʯ • σʔλͷऔಘ • ୳ࡧతσʔλੳͱલॲཧ •
༧ଌϞσϧͷߏங • ݁ՌͷղऍͱϏδωε׆༻
ࣗݾհ • ੴݪↅଠ !VQVSB • ຊܦࡁ৽ฉࣾ ݄ೖࣾ • σʔλΞφϦετˍΤϯδχΞ •
େֶ࣌ɿֶֶ෦ɺ՝֎׆ಈେֶ৽ฉ • झຯɿ,BHHMFɺڝϓϩɺϒϩά ʢ݄BEWFOUDBMFOEBSͳͲͰຊࣥචʣ
σʔλυϦϒϯνʔϜ • αʔϏεاըɾ։ൃӦۀɾϚʔέςΟϯάͰ ʮσʔλΛۙʹʯ • ୯ͳΔੳ͚ͩͰͳ͘ɺج൫ͷඋɺଌఆ߲ͷ ઃܭɺۀޮԽʹ͚ͨڥඋͳͲ • ར༻ݴޠɿ42- 1ZUIPO
3 /PEFKT ຊޠ
ຊͷ • ࣗݾհ • σʔλಓͱʮܦిࢠ൛1SPʯ • σʔλͷऔಘ • ୳ࡧతσʔλੳͱલॲཧ •
༧ଌϞσϧͷߏங • ݁ՌͷղऍͱϏδωε׆༻
σʔλಓ • σʔλυϦϒϯΛՃ͢Δڭҭ੍ʢʙʣ • ੳ୲ऀ͚ͩͰͳ͘ฤूɾӦۀɾࠂͷؔऀΒ ͕ɺ42-σʔλʹجͮ͘1%$"ͷճ͠ํΛֶͿ • Χ݄ʹΘͨΓिʹҰɺۀ࣌ؒͷ࣌ؒ ͷͰूதతʹऔΓΉ
ۀͷݹ͍ձࣾͰσʔλͷຽओԽΛਐΊͨ IUUQTTQFBLFSEFDLDPNZPTVLFTV[VLJOJLLFJEBUBESJWFO
ػցֶशτϨʔχϯά • σʔλಓͷൃల൛ • ֎෦ߨࢣট͖ɺػցֶशͷཧϏδωεԠ༻ ͢ΔͨΊͷϊϋͳͲΛֶͿ • ύοέʔδΛΘͳ͍ػցֶशΞϧΰϦζϜͷ࣮ ͔Β࢝Ίɺ࠷ऴతʹػցֶशΛ༻͍ͯࣗࣾαʔϏε ͷσʔλΛੳ
ܦిࢠ൛1SP • ๏ਓ͚ͷʮܦిࢠ൛ʯ IUUQTQSOJLLFJDPNQSP • ෳਓͰهࣄͷίϝϯτڞ༗͕Ͱ͖Δάϧʔϓ ػೳͳͲɺݸਓܖͷܦిࢠ൛ʹͳ͍ػೳɾ ίϯςϯπ͕ॆ࣮ • ຊܖલͷແྉτϥΠΞϧΛఏڙ
• ແྉτϥΠΞϧ͔ΒຊܖʹࢸΔׂ߹ɺ͢ͳΘͪ ʮຊܖʯɺച্ʹ݁͢Δॏཁͳࢦඪ
ࠓճͷੳͷత • ຊܖͷ্Λࢦ͠ɺաڈʹແྉτϥΠΞϧ ͔Βຊܖͨ͠ʗ͠ͳ͔ͬͨϢʔβΛରʹ͠ɺ ͦΕͧΕͲͷΑ͏ͳಛ͕͋Δ͔Ѳ • Ϣʔβͷଐੑใར༻ʹؔ͢Δใ͔Βɺ ػցֶशΛ༻͍Δ͜ͱͰେྔͷσʔλΛॲཧ͠ɺ ຊܖ͢Δ͔൱͔ʹؔΘΔಛΛఆੑతͰͳ͘ ఆྔతʹಛఆ
ಛྔͷॏཁ આ໌ม !ɿ Ϣʔβଐੑར༻ user_id "# "$ ... "%
& 00000001 0 00000002 1 00000003 0 తม yɿ ຊܖʹࢸ͔ͬͨ൱͔ ػցֶशϞσϧ ಗ໊Խ͞Εͨ*% ༧ଌʹ༻͍ͨಛͷॏཁΛࢉग़ ˠຊܖʹӨڹ͢ΔಛͱԿ͔ʁ
ຊͷ • ࣗݾհ • σʔλಓͱʮܦిࢠ൛1SPʯ • σʔλͷऔಘ • ୳ࡧతσʔλੳͱલॲཧ •
༧ଌϞσϧͷߏங • ݁ՌͷղऍͱϏδωε׆༻
"UMBT • ͨ͠ϦΞϧλΠϜσʔλॲཧج൫ʮ"UMBTʯ ϦΞϧλΠϜσʔλॲཧج൫ ʮ"UMBTʯ ͷιʔείʔυΛެ։͠·͢ IUUQTIBDLOJLLFJDPNCMPH
[email protected]
@QSPKFDU
42- 1ZUIPOͰੳ • 3FEBTI্Ͱ42-Λॻ͖ɺσʔλΛऔಘ • ࠓճػցֶशΛ༻͍ͨൺֱతෳࡶͳੳΛߦ͏ ߹্ɺ42-Ͱσʔλऔಘ·ͰΛѻ͍ɺΓͷ ॲཧ1ZUIPOΛར༻ • ˞,JCBOB
%0.0 34UVEJPͳͲར༻Ͱ͖Δ
ຊͷ • ࣗݾհ • σʔλಓͱʮܦిࢠ൛1SPʯ • σʔλͷऔಘ • ୳ࡧతσʔλੳͱલॲཧ •
༧ଌϞσϧͷߏங • ݁ՌͷղऍͱϏδωε׆༻
୳ࡧతσʔλੳʢ&%"ʣ • औಘͨ͠σʔλͷ֤ಛͷɺܽམͷ༗ແ ͳͲΛ֬ೝ • ݸਓతͳݟղͱͯ͠ɺϏδωεͷੈքͰσʔλΛ ѻ্͍ͬͯ͘Ͱಛʹॏཁͳաఔ • ,BHHMFͳͲͱൺɺϏδωεͰղܾ͖͢ Λಛఆ͠ԾઆΛཱͯΔ͜ͱʹՁ͕͋Δ
σʔλΛදࣔ͢Δ • ଐੑใ͕ఔɺΞΫηεใ͕ఔ
σʔλͷ֓ཁΛ͔ͭΉ • جૅ౷ܭྔܽଛΛோΊΔ • ! == 0 ͕ଟ͍ෆۉߧσʔλ • ʮอଘهࣄʯʮࣗ༝ճͷଐੑใʯʹܽଛ
• ˞લऀ42-ͷॻ͖ํͷʢKPJOʣ
U4/&ͰՄࢹԽ • ߴ࣍ݩσʔλͷ࣍ݩݮͷख๏ • ԫ৭ͷ ! == 1 ͕ൺֱత·ͱ·ͬͨҐஔʹ
ܽଛΧςΰϦมͷॲཧ • ܽଛ͕ଟ͗͢Δมআ • ʮอଘهࣄʯͷܽଛͰຒΊΔ • ΧςΰϦมμϛʔมʹ
-FBLBHFͷআ • ༧ଌͷରͱͳΔʹؔ͢Δ༧ظͤ͵ใֶ͕श σʔλʹଘࡏ͢ΔͨΊɺػցֶशΞϧΰϦζϜ ͕ඇݱ࣮తʹߴ͍ਫ਼Λࣔ͢ݱ • ࠓճʮຊܖਃ͠ࠐΈखଓ͖ϖʔδͷӾཡʯ ͕-FBLBHFʹ • ຊܖΛਃ͠ࠐΉखଓ͖ϖʔδΛӾཡ͍ͯ͠Δ
Ϣʔβɺવ΄΅ͷ֬ͰຊܖʹࢸΔ
ຊͷ • ࣗݾհ • σʔλಓͱʮܦిࢠ൛1SPʯ • σʔλͷऔಘ • ୳ࡧతσʔλੳͱલॲཧ •
༧ଌϞσϧͷߏங • ݁ՌͷղऍͱϏδωε׆༻
ػցֶशϞσϧͷબఆ • ਖ਼ղ"6$ͰϞσϧͷਫ਼Λൺֱ
(SBEJFOU#PPTUJOH$MBTTJGJFS • TLMFBSOͷޯϒʔεςΟϯάܾఆΛ࠾༻ • ཧ༝ᶃ ಛͷॏཁΛࢉग़Ͱ͖ɺతʹ߹க • ཧ༝ᶄ 47$ͱൺೋྨҎ֎ʹԠ༻͍͢͠ •
(SJE4FBSDI$7ͰϋΠύʔύϥϝʔλௐ • ަࠩݕূͷ"6$Ͱఔ
ຊͷ • ࣗݾհ • σʔλಓͱʮܦిࢠ൛1SPʯ • σʔλͷऔಘ • ୳ࡧతσʔλੳͱલॲཧ •
༧ଌϞσϧͷߏங • ݁ՌͷղऍͱϏδωε׆༻
ಛͷॏཁ • ࠓճͷ༧ଌϞσϧʹ͓͚ΔಛͷॏཁΛग़ྗ • ˞աʹಛͷॏཁΛ৴ͣ͡ɺཧతഎܠΛҙࣝ ͯ͠৻ॏʹղऍ͢Δඞཁ͕͋Δ • αʔϏεӦۀɾϚʔέςΟϯάͷ୲ऀʹڞ༗ ͠ɺࠓޙͷࢪࡦʹ͚ͨٞͷࡐྉʹ
·ͱΊ • ػցֶशΛ༻͍ͯܦిࢠ൛1SPͷϢʔβੳΛ ࣮ࢪ͠ɺແྉτϥΠΞϧ͔ΒຊܖʹࢸΔཁҼͱ ͳΔಛΛఆྔతʹಛఆͨ͠ • Ұݟʮݹष͍ʯຊܦࡁ৽ฉࣾͰɺσʔλ׆༻͕ ੵۃతʹల։͞Ε͍ͯΔ ʢσʔλಓɾσʔλج൫ɾػցֶशͳͲʣ