Upgrade to Pro
— share decks privately, control downloads, hide ads and more …
Speaker Deck
Features
Speaker Deck
PRO
Sign in
Sign up for free
Search
Search
特徴量の重要度計算を実装する ランダムフォレストを実装する
Search
NearMeの技術発表資料です
PRO
September 28, 2023
170
0
Share
Embed
Copy iframe code
Copy JS code
Copy link
Start on current slide
特徴量の重要度計算を実装する ランダムフォレストを実装する
NearMeの技術発表資料です
PRO
September 28, 2023
More Decks by NearMeの技術発表資料です
See All by NearMeの技術発表資料です
Claude Code × git worktree で並列開発 (続き) -差分のサービスだけを併設する-
nearme_tech
PRO
1
50
Claude Code × git worktree で並列開発 — サブモジュール構成のリポジトリで成立させる —
nearme_tech
PRO
0
54
LLM + 強化学習
nearme_tech
PRO
0
31
PosthogのA/Bテスト機能の紹介
nearme_tech
PRO
1
45
AIフレンドリーなプロダクトに向けて
nearme_tech
PRO
2
65
初めてのLean言語
nearme_tech
PRO
0
100
Apache Airflow Workflow orchestration without turning cron into spaghetti
nearme_tech
PRO
2
37
実務で役立つ幾何学 ボロノイ図の基礎から グラフ・ネットワーク応用まで
nearme_tech
PRO
1
76
SQL/ID抽出タスクから考える 実践的なハルシネーション対策
nearme_tech
PRO
1
88
Featured
See All Featured
Google's AI Overviews - The New Search
badams
0
1.6k
Automating Front-end Workflow
addyosmani
1369
210k
Why Your Marketing Sucks and What You Can Do About It - Sophie Logan
marketingsoph
0
400
Building Experiences: Design Systems, User Experience, and Full Site Editing
marktimemedia
0
600
Digital Ethics as a Driver of Design Innovation
axbom
PRO
1
410
Put a Button on it: Removing Barriers to Going Fast.
kastner
60
4.6k
Visual Storytelling: How to be a Superhuman Communicator
reverentgeek
2
650
Building an army of robots
kneath
306
46k
Refactoring Trust on Your Teams (GOTO; Chicago 2020)
rmw
35
3.8k
Intergalactic Javascript Robots from Outer Space
tanoku
273
27k
No one is an island. Learnings from fostering a developers community.
thoeni
21
3.8k
Helping Users Find Their Own Way: Creating Modern Search Experiences
danielanewman
31
3.4k
Transcript
0 特徴量の重要度計算を実装する ランダムフォレストを実装する 2023-09-15 第60回NearMe技術勉強会 Takuma Kakinoue
1 今回の実装のコミット https://github.com/kakky-hacker/algorithm_sandbox/commit/8f466be9c1fae0d7 bf3bc1bff1a5d1327e4a324d 1
2 Feature importanceの実装 • Feature importance の2種類の定義 • 分割に使われた回数(split) •
分割において得られた情報ゲインの総和(gain) ← こちらを採⽤ class Node: … def feature_importance(self, importance_value): if not self.is_leaf: importance_value[self.split_feature_index] += self.gain self.left_node.feature_importance(importance_value) self.right_node.feature_importance(importance_value) return importance_value ルートノードから 情報ゲインを再帰的に ⾜し合わせていく
3 ランダムフォレストの実装 • 実装としては決定⽊を並列にアンサンブルするだけ • 既存の決定⽊の変更箇所(core.py) ◦ 決定⽊で使⽤しない特徴量をマスク形式で⼊⼒に加える • 新規実装(random_forest.py)
◦ ランダムフォレストのメイン処理 ▪ 特徴量マスクをランダムに⽣成して決定⽊学習 ▪ 推論時は、各決定⽊の出⼒を多数決 3
4 Borutaの実装(途中) • Borutaとは • 特徴量選択⼿法の1つ • ランダムな特徴量(shadow feature)を⼊⼒に結合し、 オリジナル特徴量とランダムな特徴量の重要度を⽐較する
• 学習を100回⾏い、上記⽐較のt検定を実施し、 真に重要なオリジナル特徴量を選択する 4
5 次回予告 • ⾃作ランダムフォレストとscikit-learnのランダムフォレストの 性能を⽐較する。 • 勾配ブースティングを実装する。 • ⾃動ハイパーパラメータチューニングを実装する。 5
6 Thank you