Upgrade to Pro
— share decks privately, control downloads, hide ads and more …
Speaker Deck
Sign up for free
Menu
Search
Features
All features
Private URLs
Password Protection
Custom URLS
Scheduled publishing
Remove Branding
Restrict embedding
Deck Collections
Notes
Features
All features
Private URLs
Password Protection
Custom URLS
Scheduled publishing
Remove Branding
Restrict embedding
Deck Collections
Notes
Explore
Featured decks
Featured speakers
Programming
Technology
Storyboards
Explore
Featured decks
Featured speakers
Programming
Technology
Storyboards
Pricing
Search
Sign in
Sign up for free
atmacup8 振り返り会登壇資料
Search
Sponsored
·
Your Podcast. Everywhere. Effortlessly.
Share. Educate. Inspire. Entertain. You do you. We'll handle the rest.
→
Sekine Hiroto
February 25, 2021
Science
450
0
Share
Embed
Copy iframe code
Copy JS code
Copy link
Start on current slide
atmacup8 振り返り会登壇資料
Sekine Hiroto
February 25, 2021
More Decks by Sekine Hiroto
See All by Sekine Hiroto
推薦システムと多目的最適化
hiroto0227
2
2.7k
WWW2020 online報告会
hiroto0227
0
680
Other Decks in Science
See All in Science
因果推論と機械学習
sshimizu2006
1
1.4k
Conversation is the New Dashboard: 属人性を排除する第4世代BIツールの勢力図
shomaekawa
1
650
2026 Introduction to University Math 01
kanaya
0
140
データベース12: 正規化(2/2) - データ従属性に基づく正規化
trycycle
PRO
0
1.7k
Does the Efficient Compute Frontier Represent New Physics?
drqz
0
120
Toward Causal Scientific Discovery with AI
sshimizu2006
0
180
[TMLR 2026, Featured Certification] Double Bounded α-Divergence Optimization for Density Estimation
gkazunii
1
110
Understanding CVP Waveforms: Interpretation and Clinical Implications in Anesthesiology
taka88
0
860
人生を変えた一冊「独学大全」のはなし / Self-study ENCYCLOPEDIA: The Book Which Change My Life #独学大全 #EM推し本
expajp
0
210
AI bij literatuuronderzoek in de wetenschap
voginip
0
250
20260722【JAWS-UG東京 ランチタイムLT会 #37④】AWS Well-Architectedフレームワークに沿った回答をするAIエージェントを作ってみた
nozakijcom
1
140
[第67回 CV勉強会@関東] CV × Scientific Figures / kantoCV 67th CVPR 2026
lychee1223
0
200
Featured
See All Featured
More Than Pixels: Becoming A User Experience Designer
marktimemedia
3
520
A Guide to Academic Writing Using Generative AI - A Workshop
ks91
PRO
1
440
Avoiding the “Bad Training, Faster” Trap in the Age of AI
tmiket
0
230
Hiding What from Whom? A Critical Review of the History of Programming languages for Music
tomoyanonymous
3
1.2k
Unlocking the hidden potential of vector embeddings in international SEO
frankvandijk
0
930
A brief & incomplete history of UX Design for the World Wide Web: 1989–2019
jct
2
500
Stewardship and Sustainability of Urban and Community Forests
pwiseman
0
520
The Power of CSS Pseudo Elements
geoffreycrofte
82
6.6k
The B2B funnel & how to create a winning content strategy
katarinadahlin
PRO
1
510
How People are Using Generative and Agentic AI to Supercharge Their Products, Projects, Services and Value Streams Today
helenjbeal
1
310
It's Worth the Effort
3n
188
29k
How to make the Groovebox
asonas
2
2.4k
Transcript
atmacup 8 振り返り会 Sekine Hiroto
• Sekine Hiroto • twitter: @ndnto • github: @hiroto0227 •
大学では自然言語処理を研究 • 20卒でWantedly, inc.にJoin。(推薦やデータサイエンス) 趣味 • ビール (特に海外ビール) 自己紹介
目次 • 今回のコンペの進め方 • 作成した特徴量 • 今回できなかったこと
• データ分析コンペを始めるのには最適! • 1週間という短い期間で集中して取り組む! • ディスカッションに知りたかったことやアイディアが丁寧に書かれている! atmacupへの印象 今回も多くのディスカッションを参考にさせていただきました - RMSLEを最適化する小技
- シリーズ名ごとにGroup KFold
進め方 • @hakubishin3 と @yu-ya4 とチームを組んで進めた。 ◦ チームを組んだ目的としては、 @hakubishin3 から知見を共有してもらう。
◦ 初日の夜に効きそうな特徴量を聞いて、そこからは個人でガンガン。 • まず始めにSubmit! ◦ 初心者にとってはここが難関 -> 1日集中してサブミットまでできる仕組みを作成。 ◦ Data Load, Preprocessor, Model, Training, Predict, Submit • 特徴量を自由に追加できる形にしておく。 ◦ 1つの関数で1つの処理を行う。 ◦ 関数の入力、出力を合わせておくことで、 for文で回せるようにする。 • そこからは、アイディアと面倒くさがらずにできるかの勝負! ◦ 効きそうな特徴量から作っていく。
作成した特徴量 • Aggregation Feature ◦ カテゴリ変数(PublisherやDeveloper, Nameを含む)ごとに、Year_of_ReleaseやCritic_Scoreなど に対し集計処理 • Diff
Feature ◦ Aggregation Featureの集計した平均値と各レコードの平均値の差をとる。 • Target Encoding ◦ カテゴリ変数ごとに、 xx_Salesに対し集計処理 ◦ DeveloperやPublisherを入れたらリークした。 • LDAによる分散表現 ◦ 分散表現にすることで、 Aggregation Featureでは与えられないような角度からの特徴を得たい。 • Rank Feature
PublisherとDeveloperについて • TrainとTestの分け方がPublisherによって分割されていると想定 • Publisherを直接使用したところ、CVの値は劇的に上がったが、LBの値が下がると いう現象が起きた。 • 「Nintendo」というPublisherを使用できる特徴量で表現したい。 ◦ UserScoreの平均が高く、幅広いジャンルのゲームを作成している
◦ ここでGlobal_Salesが多いという情報を使用してしまうと、それを TestデータのPublisherで再 現することはできない。 • Publisherごとに特徴量の平均、分散、パーセンタイル値、Countなどの集計値をと る。
データをざっと見る。 • PlatformやGenreやPlatformはユニーク数も少なくTrainとTestに満遍なく入ってい たため、そのまま使用することができた。 ◦ Target EncodingもPlatformごとやGenreごとなどで使用できた。 • Critic_ScoreやUser_Scoreなどは50%程度欠損値である。 •
DeveloperやRatingについても欠損値が多かった。 • このNull値を補えるように特徴量を作成したい。
力を入れたところ (Series Nameの名寄せ) • どこまでをシリーズとみなすかがゲームによって異なる。 ◦ LEGO Batman を LEGOとするか?
LEGO Batmanとするか? 方法 • Nameの最初から5gramを見る。 ◦ 3回以上出現した5gramがあれば、辞書に追加 • Nameの最初から4gramを見る。 ... • Nameと辞書にマッチするもののうち、最も単語数の 多いものをそのシリーズ名とする。
RankFeature • 各カテゴリごとにYear_of_Releaseを並べたもの • AggregateFeatureばかりを作成していたため、同じカテゴリのものは同じ特徴量と なってしまうことを変えたかった。 • Year_of_Releaseが同じ場合は登場順でランク付けを行う。 • Genre_Year_of_Releaseのカテゴリでランク付けを行った。
• リークにつながり、Public Scoreが0.78から0.69に上がった。 最も効いた特徴量
None
Feature Importance
今後に向けて必要だと感じたところ • Preporcessingのクラス ◦ 特徴量が多くなってくると、各特徴量の依存関係の整合性が合わなくなりそう。 • 特徴量のNaming ◦ Aggregation Featureなどは、何のカラムを
Group Byして、何のカラムに対しての平均なのか?と いうのをルールとして決めておくことで、理解度や Namingの迷いがなくなる。 • 何の特徴量がなぜ効いてるか? ◦ 初めは特徴量を作成して学習させて、後からなぜそれが効いたのか、効かなかったかを考えようと したが、後半はスコアを伸ばしたい気持ちが強くて、なぜ効いたかを考えられなかった。 • 特徴量作成のネタ切れ
ありがとうございました! 楽しかったです!!