Upgrade to Pro
— share decks privately, control downloads, hide ads and more …
Speaker Deck
Sign up for free
Menu
Search
Features
All features
Private URLs
Password Protection
Custom URLS
Scheduled publishing
Remove Branding
Restrict embedding
Deck Collections
Notes
Features
All features
Private URLs
Password Protection
Custom URLS
Scheduled publishing
Remove Branding
Restrict embedding
Deck Collections
Notes
Explore
Featured decks
Featured speakers
Programming
Technology
Storyboards
Explore
Featured decks
Featured speakers
Programming
Technology
Storyboards
Pricing
Search
Sign in
Sign up for free
Text Mining: Exploratory Data Analysis to Machi...
Search
Sponsored
·
SiteGround - Reliable hosting with speed, security, and support you can count on.
→
Julia Silge
March 04, 2019
Technology
280
1
Share
Embed
Copy iframe code
Copy JS code
Copy link
Start on current slide
Text Mining: Exploratory Data Analysis to Machine Learning
March 2019 talk at WiDS Salt Lake City regional event
Julia Silge
March 04, 2019
More Decks by Julia Silge
See All by Julia Silge
Introducing Positron
juliasilge
1
420
The right tool for the job
juliasilge
0
100
Good practices for applied machine learning
juliasilge
0
280
Applied machine learning with tidymodels
juliasilge
0
190
Maintaining an R Package
juliasilge
0
470
Publishing the Stack Overflow Developer Survey
juliasilge
2
120
Text Mining Using Tidy Data Principles
juliasilge
0
210
North American Developer Hiring Landscape
juliasilge
0
100
Understanding Principal Component Analysis Using Stack Overflow Data
juliasilge
13
4.6k
Other Decks in Technology
See All in Technology
BedrockとLambdaで作る リアルタイム進行型推理ゲーム
kawametho
0
190
20260930_Gemma4_Hands-on
tsho
0
230
AI Made Us Faster at Solving the Wrong Problems
marceloancelmo
0
180
私の推しは「聞いてから進む」AIです -AI-DLCに一人でアプリを作らせた話
yama3133
1
180
継続的なAIコスト最適化
bgpat
1
180
DWH の限界突破! AI エージェント向け爆速リアルタイムデータ基盤 ClickHouse
jozono
1
210
1万名の社員が使う認証基盤で どう信頼性を担保するか?
kairim0
0
210
SREでアラート疲れを 解決しよう!
kairim0
1
210
しろおびから始める!ログ活用の第一歩
shumei_ito
0
290
形式手法を使って仕様をコーディングしよう
mikanichinose
0
180
検証フェーズはAIの回答を判断するための学習機会
toru_kubota
2
510
組み立てて楽しむ AWS Blocks 入門
kmiya84377
0
220
Featured
See All Featured
Exploring the Power of Turbo Streams & Action Cable | RailsConf2023
kevinliebholz
37
6.6k
Optimizing for Happiness
mojombo
378
71k
For a Future-Friendly Web
brad_frost
184
10k
The Curse of the Amulet
leimatthew05
3
15k
Ten Tips & Tricks for a 🌱 transition
stuffmc
1
250
職位にかかわらず全員がリーダーシップを発揮するチーム作り / Building a team where everyone can demonstrate leadership regardless of position
madoxten
69
66k
The Myth of the Modular Monolith - Day 2 Keynote - Rails World 2024
eileencodes
28
3.7k
Collaborative Software Design: How to facilitate domain modelling decisions
baasie
1
340
Avoiding the “Bad Training, Faster” Trap in the Age of AI
tmiket
0
260
Measuring & Analyzing Core Web Vitals
bluesmoon
9
1k
Leveraging Curiosity to Care for An Aging Population
cassininazir
1
530
Into the Great Unknown - MozCon
thekraken
41
2.8k
Transcript
T E X T M I N I N G
EXPLORATORY DATA ANALYSIS TO MACHINE LEARNING
HELLO T I D Y T E X T Data
Scientist at Stack Overflow @juliasilge https://juliasilge.com/ I’m Julia Silge
T I D Y T E X T TEXT DATA
IS INCREASINGLY IMPORTANT
T I D Y T E X T TEXT DATA
IS INCREASINGLY IMPORTANT NLP TRAINING IS SCARCE ON THE GROUND
TIDY DATA PRINCIPLES + COUNT-BASED METHODS = T I D
Y T E X T
https://github.com/juliasilge/tidytext
https://github.com/juliasilge/tidytext
http://tidytextmining.com/
T I D Y T E X T EXPLORATORY DATA
ANALYSIS N-GRAMS AND MORE WORDS MACHINE LEARNING
EXPLORATORY DATA ANALYSIS T I D Y T E X
T
from the Washington Post’s Wonkblog
from the Washington Post’s Wonkblog
D3 visualization on Glitch
WHAT IS A DOCUMENT ABOUT? T I D Y T
E X T TERM FREQUENCY INVERSE DOCUMENT FREQUENCY
None
None
• As part of the NASA Datanauts program, I worked
on a project to understand NASA datasets • Metadata includes title, description, keywords, etc
None
T A K I N G T I D Y
T E X T T O T H E N E X T L E V E L N-GRAMS, NETWORKS, & NEGATION
None
None
None
None
None
T A K I N G T I D Y
T E X T T O T H E N E X T L E V E L TOPIC MODELING
TOPIC MODELING T I D Y T E X T
•Each DOCUMENT = mixture of topics •Each TOPIC = mixture of words
None
None
None
None
T A K I N G T I D Y
T E X T T O T H E N E X T L E V E L TEXT CLASSIFICATION
TRAIN A GLMNET MODEL T I D Y T E
X T
TEXT CLASSIFICATION T I D Y T E X T
> library(glmnet) > library(doMC) > registerDoMC(cores = 8) > > is_jane <- books_joined$title == "Pride and Prejudice" > > model <- cv.glmnet(sparse_words, is_jane, family = "binomial", + parallel = TRUE, keep = TRUE)
None
None
THANK YOU T I D Y T E X T
@juliasilge https://juliasilge.com JULIA SILGE
THANK YOU T I D Y T E X T
@juliasilge https://juliasilge.com Author portraits from Wikimedia Photos by Glen Noble and Kimberly Farmer on Unsplash JULIA SILGE