Upgrade to Pro
— share decks privately, control downloads, hide ads and more …
Speaker Deck
Features
Speaker Deck
PRO
Sign in
Sign up for free
Search
Search
Pyspark - produtividade e poder de processamento
Search
Sponsored
·
Ship Features Fearlessly
Turn features on and off without deploys. Used by thousands of Ruby developers.
→
Felipe cruz
November 10, 2015
Technology
77
1
Share
Embed
Copy iframe code
Copy JS code
Copy link
Start on current slide
Pyspark - produtividade e poder de processamento
Felipe cruz
November 10, 2015
More Decks by Felipe cruz
See All by Felipe cruz
Recomendação - Algoritmos de Filtragem Colaborativa
felipecruz
0
450
Coleta Massiva de Dados
felipecruz
2
140
TDC 2014 - Machine Learning Guerrilha
felipecruz
0
300
Python & C - Formas de Integração
felipecruz
0
140
Other Decks in Technology
See All in Technology
まちスペース®とデジタルツインと「まちづくり」
hiro_ogi
0
120
Sets in Go
ramalho
1
620
Bits Agent Builder の⼊⾨と活⽤事例
nulabinc
PRO
0
140
ブラウザ研修 2026
recruitengineers
PRO
6
990
【CEDEC2026】『Relink』を拡張せよ - 『GRANBLUE FANTASY: Relink - Endless Ragnarok』の開発速度と品質を守るCI運用
cygames
PRO
0
170
侵入は突然に 〜 IoTマルウェアと悪用される家庭の機器 ~ / When Intrusion Strikes: IoT Malware and the Abuse of Home Devices
nttcom
0
3.7k
LanceDB入門
mocobeta
7
500
Master Dataグループ紹介資料
sansan33
PRO
1
4.8k
猫付きpingコマンドを自作
uyuki234
0
230
PLATEAU で バーチャル花火大会
tatsuya1970
0
150
老害フォレンジッカーはAI羊の夢を見るか?
tadmaddad
0
330
つくって納得、つかって実感! 大規模言語モデルことはじめ ver2.0
recruitengineers
PRO
4
1.6k
Featured
See All Featured
ピンチをチャンスに:未来をつくるプロダクトロードマップ #pmconf2020
aki_iinuma
128
56k
How to optimise 3,500 product descriptions for ecommerce in one day using ChatGPT
katarinadahlin
PRO
2
3.8k
Impact Scores and Hybrid Strategies: The future of link building
tamaranovitovic
0
390
Crafting Experiences
bethany
1
250
Jamie Indigo - Trashchat’s Guide to Black Boxes: Technical SEO Tactics for LLMs
techseoconnect
PRO
0
620
Have SEOs Ruined the Internet? - User Awareness of SEO in 2025
akashhashmi
0
410
Taking LLMs out of the black box: A practical guide to human-in-the-loop distillation
inesmontani
PRO
3
2.3k
Heart Work Chapter 1 - Part 1
lfama
PRO
8
36k
Building Experiences: Design Systems, User Experience, and Full Site Editing
marktimemedia
0
570
Bootstrapping a Software Product
garrettdimon
PRO
307
120k
The Hidden Cost of Media on the Web [PixelPalooza 2025]
tammyeverts
2
460
The Illustrated Guide to Node.js - THAT Conference 2024
reverentgeek
1
430
Transcript
PySpark PySpark Produtividade e poder de processamento
Quem? Quem? github.com/felipecruz github.com/felipecruz @ @felipe felipej jcruz cruz
Agenda Agenda Map-Reduce Pyspark
Maior oferta feita em Maior oferta feita em uma semana
na uma semana na BOVESPA? BOVESPA?
Motivação Motivação highest_offer = max(offers)
None
None
? ?
? ?
highest_offers_1 = max(offers_partition1) highest_offers_2 = max(offers_partition2)
highest_offers_1 = max(offers_partition1) highest_offers_2 = max(offers_partition2) max(highest_offers_1, highest_offers_2)
calma... calma...
Map
Map Reduce
Map-Reduce Map-Reduce não é divisão e conquista (que pode ser
implementada com map-reduce)
Aplicações Aplicações Filtragem Distintos Top K Por valor Sumarização Índice
invertido Contagem de palavras Estruturação Ordenação Particionamento Embaralhamento Join Inner join Produto cartesiano nosso exemplo K = 1
PySpark PySpark
Funcionalidades centrais Funcionalidades centrais Map-Reduce RDD, DataFrames & SQL MLlib
Streaming GraphX
Map Map & Reduce & Reduce >>> prices = sc.textFile('s3n://prognoos-pyspark/*.gz')
\ ... .filter(lambda x: x.count(';') > 14) \ ... .map(lambda x: [s.strip() for s in x.split(';')]) \ ... .map(lambda x: (x[1], x[8], x[15])) ... >>> prices.take(2) [(u'ABEVA70', u'000000000000.350000', u'000000000000008300'), (u'ABEVA70', u'000000000000.350000', u'000000000000007100')]
Map & Map & Reduce Reduce >>> prices = sc.textFile('ftp://*.gz')
\ ... .filter(lambda x: x.count(';') > 14) \ ... .map(lambda x: [s.strip() for s in x.split(';')]) \ ... .map(lambda x: (x[1], float(x[8]), x[15])) ... >>> sum_all = prices.map(lambda x: x[2])\ ... .reduce(lambda x, y: x + y) ... >>> sum_all 1532623750.0
from datetime import datetime strpt = lambda x: datetime.strptime(x, '%H:%M:%S.%f')
f = float negs = sc.textFile('s3n://prognoos-pyspark/NEG/*.gz') \ .filter(lambda x: x.count(';') > 14) \ .map(lambda x: [s.strip() for s in x.split(';')]) \ .map(lambda x: (strpt(x[5]), 'NEG', x[1], f(x[3]), f(x[16]), x[17])) buys = sc.textFile('s3n://prognoos-pyspark/CPA/*.gz') \ .filter(lambda x: x.count(';') > 14) \ .map(lambda x: [s.strip() for s in x.split(';')]) \ .map(lambda x: (strpt(x[6]), 'CPA', x[1], f(x[8]), x[15], None)) sell = sc.textFile('s3n://prognoos-pyspark/VDA/*.gz') \ .filter(lambda x: x.count(';') > 14) \ .map(lambda x: [s.strip() for s in x.split(';')]) \ .map(lambda x: (strpt(x[6]), 'VDA', x[1], f(x[8]), None, x[15])) all_operations = negs.union(buys).union(sell) total = all_operations.count() # total = 52980676
... nem tudo são ... nem tudo são flores flores
data = sc.parallelize(['aa', 'bb', 'ab', 'bc']) def _filter(data): sts =
['a', 'b'] rets = [] for st in sts: rets.append((st, data.filter(lambda x: x.startswith(st)))) return rets rdds = _filter(data) for st, rdd in rdds: print((st, rdd.collect())) # ('a', ['bb', 'bc']) # ('b', ['bb', 'bc']) Python - Anti-pattern - não faça!!
DataFrames & SQL DataFrames & SQL
DataFrame DataFrame A distributed collection of data grouped into named
columns http://spark.apache.org/docs/latest/api/python/pyspark.sql.html#pyspark.sql.DataFrame
events = negs.union(buys).union(sell).toDF() # API de DataFrame total = events.count()
# Salva pra uso posterior events.write.save('s3n://prognoos/events/', format='parquet', mode='Overwrite')
SparkSQL SparkSQL http://spark.apache.org/docs/latest/api/python/pyspark.sql.html >>> path = 's3n://prognoos-test/events' >>> table_name =
'bovespa_events' >>> events = sqlContext.read.parquet(path) >>> events.registerTempTable(table_name) >>> total_events = sqlContext.sql(''' select count(*) from bovespa_events ''')
Spark em produção Spark em produção Standalone Hadoop/Yarn Mesos
Spark em produção Spark em produção
Dúvidas? Dúvidas? @felipejcruz @felipejcruz github.com/felipecruz github.com/felipecruz