Upgrade to Pro
— share decks privately, control downloads, hide ads and more …
Speaker Deck
Features
Speaker Deck
PRO
Sign in
Sign up for free
Search
Search
Spark Machine Learning 101 @HadoopCon
Search
Chu-Yu Hsu
September 19, 2015
Technology
420
1
Share
Embed
Copy iframe code
Copy JS code
Copy link
Start on current slide
Spark Machine Learning 101 @HadoopCon
Chu-Yu Hsu
September 19, 2015
Other Decks in Technology
See All in Technology
Webの技術とガジェットで子どもも大人も楽しめるワクワク体験を提供する / Qiita Tech Festa Day 2026
you
PRO
1
310
AIエージェントがあれば技術書なんてすぐ書けるでしょ→無理でした
watany
6
1.1k
データベース研修【MIXI 26新卒技術研修】
mixi_engineers
PRO
1
670
検索技術知識0のエンジニアが広告検索システムを内製化して運用するまで
lycorptech_jp
PRO
0
150
AI研修(Day2)【MIXI 26新卒技術研修】
mixi_engineers
PRO
2
1.5k
2026年のソフトウェア開発を考える(2026/07版) / Agentic Software Engineering 2026-07 Findy Edition
twada
PRO
31
21k
システム監視入門
grimoh
5
750
WEBフロントエンド研修【MIXI 26新卒技術研修】
mixi_engineers
PRO
1
730
なぜ、あなたのAPIは使われないのか? AX時代の設計原則、ガードレール、運用体制
yokawasa
1
260
大 AI 時代におけるC# の事情 ~ぶっちゃけトークを交えながら~
nenonaninu
1
540
LangfuseによるLLMOps基盤の構築と活用事例
zozotech
PRO
1
160
QA・ソフトウェアテスト研修【MIXI 26新卒技術研修】
mixi_engineers
PRO
3
1.7k
Featured
See All Featured
KATA
mclloyd
PRO
35
15k
The Illustrated Guide to Node.js - THAT Conference 2024
reverentgeek
1
420
Keith and Marios Guide to Fast Websites
keithpitt
413
23k
Navigating Weather and Climate Data
rabernat
0
420
How to Ace a Technical Interview
jacobian
281
24k
How to audit for AI Accessibility on your Front & Back End
davetheseo
0
480
Connecting the Dots Between Site Speed, User Experience & Your Business [WebExpo 2025]
tammyeverts
11
980
Building Flexible Design Systems
yeseniaperezcruz
330
40k
No one is an island. Learnings from fostering a developers community.
thoeni
21
3.8k
Visualization
eitanlees
152
17k
Refactoring Trust on Your Teams (GOTO; Chicago 2020)
rmw
35
3.7k
Product Roadmaps are Hard
iamctodd
55
12k
Transcript
Spark Machine Learning 101 Chu-Yu Hsu @ HadoopCon 2015
About Me Chu-Yu Hsu, 許儲⽻羽 • Software Engineer • Machine
Learning Practicer • Used Spark ML and Python in daily work and Kaggle competition • http://blog.chuyuhsu.ml
Outline • Introduction to Spark ML • Alternative Least Squares
(ALS) • Hands-on example
None
Apache Spark MLlib • To Make practical machine learning easy
and scalable • spark.mllib - the primary API • spark.ml - a higher-level API for constructing ML workflows Apache Spark spark.mllib spark.ml
What’s in MLlib Utilities Data types Basic statistics Classification and
regression SVM Logistic regression Linear regression Naive Bayes Decision trees Ensembles of trees Isotonic regression Collaborative filtering Alternating least squares (ALS) Clustering K-means Gaussian mixture Power iteration clustering Latent Dirichlet allocation Streaming k-means Dimensionality reduction SVD PCA Frequent pattern mining FP-growth Optimization Stochastic gradient descent Limited-memory BFGS https://spark.apache.org/docs/latest/mllib-guide.html
ML Workflow can be VERY complex
Types of Recommenders • Editorial and hand curated • Simple
aggregates • Tailored to individual users
Who Uses Recommenders
Approaches • Content based method • Item based method •
Model based method
Collaborative Filtering • One of mostly known “Recommendation Algorithm” •
Widely used in E-commerce application • The data size can be enormous • Need to be delivered as soon as possible
Collaborative Filtering Main idea: Find set N of other users
whose ratings are “similar” to X’s ratings
Users Preferences • This is a baby example • Users:
> 2M • Items: > 30M • Sparsity: > 2%
Low Rank Assumption • Matrix can be reduced to the
product of low rank matrixes • That is also understood as “latent factors” • We assume that the low factor can represent the hidden factors we do not know Action Romance Thriller
Low Rank Assumption Action Romance Thriller Action Romance Thriller
Matrix Factorization
• Our goal is to find P and Q such
that (Sum of Square Error): • Root Mean Square Error (RMSE)
Alternative Least Squares • Because p and q are both
unknown, the object function is not convex • If fix one of the unknowns > can be solved as a least squares problem
Amazon Reviews Dataset 35 million ratings, 6.6 million users, 2.4
million products on 16-node (m3.2xlarge) https://github.com/apache/spark/pull/3720
Resources
Resources
And More Resources • Source code examples https://github.com/apache/spark/tree/master/ examples •
Apache Spark JIRA https://issues.apache.org/jira/browse/spark
Dataset • MovieLens Dataset http://grouplens.org/datasets/movielens/ • “ratings.dat” UserID::MovieID::Rating::Timestamp • “movies.dat”
MovieID::Title::Genres
Conclusion • Spark MLlib grows fast, but still need some
time • Spark MLlib is a strong tool, if you use it right • Sharpening ML skills is first priority
Q&A Visit me on: http://blog.chuyuhsu.ml Github: http://github.com/ChuyuHsu Thanks
References • https://spark.apache.org/docs/latest/mllib-guide.html • http://www.slideshare.net/jeykottalam/mllib • http://www.slideshare.net/PetrZapletal1/mllib-and-machine-learning-on-spark • https://databricks.com/blog/2014/07/23/scalable-collaborative-filtering-with- spark-mllib.html
• https://github.com/apache/spark/pull/3720 • https://www.hakkalabs.co/articles/spark-mllib-making-practical-machine- learning-easy-and-scalable • http://www.slideshare.net/databricks/practical-machine-learning-pipelines- with-mllib