Upgrade to Pro
— share decks privately, control downloads, hide ads and more …
Speaker Deck
Sign up for free
Menu
Search
Features
All features
Private URLs
Password Protection
Custom URLS
Scheduled publishing
Remove Branding
Restrict embedding
Deck Collections
Notes
Features
All features
Private URLs
Password Protection
Custom URLS
Scheduled publishing
Remove Branding
Restrict embedding
Deck Collections
Notes
Explore
Featured decks
Featured speakers
Programming
Technology
Storyboards
Explore
Featured decks
Featured speakers
Programming
Technology
Storyboards
Pricing
Search
Sign in
Sign up for free
Productionizing Big Data - stories from the tre...
Search
Roksolana
September 14, 2023
Technology
100
0
Share
Embed
Copy iframe code
Copy JS code
Copy link
Start on current slide
Productionizing Big Data - stories from the trenches
Presented at ScalaDays 2023 (Madrid, Spain)
Roksolana
September 14, 2023
More Decks by Roksolana
See All by Roksolana
Pain of engineering management
roksolanad
1
130
Alice and the return to the world of pods and higher-order functions
roksolanad
0
220
Modern data pipelines in AdTech - life in the trenches
roksolanad
1
330
Alice and travelling back in time
roksolanad
0
210
Big Data at AdTech
roksolanad
0
380
Alice and the Mad Hatter: Predict or not to predict
roksolanad
0
240
Alice in the world of machine learning
roksolanad
0
150
Alice and the lost pod: practical guide to Kubernetes in Scala
roksolanad
1
380
Scala meets Kubernetes
roksolanad
0
550
Other Decks in Technology
See All in Technology
いちAWSエンジニアのAI活用を振り返る #devio2026 / devio osaka 2026 kawahara
masahirokawahara
1
300
【Findyテック文化祭ワークショップ】新卒エンジニア&採用担当と作る、 なりたい姿と今やるべき一歩
dip_tech
PRO
0
150
カンファレンスに参加した後の浮遊感とセルフケア
pauli
0
270
The kernel report
ennael
PRO
1
170
Codex概要
ymiya55
0
140
OpenClawでAzure DevOpsのWiki更新を自動化する - クラウドAIだけでは届かない場所へ
yutakaosada
0
140
DORA_Metrics.pdf
wagnerfusca
1
140
PQC移行の今 -- IETF からみた現在地
satokan
4
720
プラットフォームを「作る」、 チームに「入り込む」
sansantech
PRO
0
410
10年欲しかった音楽管理アプリを、AIと一緒に作りはじめた
judau
1
200
Argo CDとAtlantisで実現するインフラ管理のセルフサービス化──小規模SREチームで支えるプラットフォーム
cassius7
0
260
人にやさしく、AIにやさしく、書き手を選ばないIaCのガードレール再考 / Rethinking IaC Guardrails for Humans and AI Alike
kohbis
5
1.9k
Featured
See All Featured
ラッコキーワード サービス紹介資料
rakko
1
5.1M
Game over? The fight for quality and originality in the time of robots
wayneb77
1
290
SERP Conf. Vienna - Web Accessibility: Optimizing for Inclusivity and SEO
sarafernandez
2
1.6k
What Being in a Rock Band Can Teach Us About Real World SEO
427marketing
0
1.1k
Keith and Marios Guide to Fast Websites
keithpitt
413
23k
4 Signs Your Business is Dying
shpigford
187
23k
How To Speak Unicorn (iThemes Webinar)
marktimemedia
1
590
Primal Persuasion: How to Engage the Brain for Learning That Lasts
tmiket
0
490
Java REST API Framework Comparison - PWX 2021
mraible
34
9.7k
Gemini Prompt Engineering: Practical Techniques for Tangible AI Outcomes
mfonobong
2
560
How To Stay Up To Date on Web Technology
chriscoyier
790
250k
JavaScript: Past, Present, and Future - NDC Porto 2020
reverentgeek
52
6.1k
Transcript
Productionizing big data - stories from the trenches
Roksolana Diachuk •Engineering manager at Captify •Women Who Code Kyiv
Data Engineering Lead •Speaker
AdTech methodologies deliver the right content at the right time
to the right consumer AdTech
None
You have your pipelines in production What’s next?
Types of issues • Low performance • Human errors •
Data source errors
Story #1. Unlucky query
Problem Drop 13 months of user profiles
Reporting
Problem 13 months hour=22042001
Loading mechanism loader.ImpalaLoaderConfig.periodToLoad: “P5D” loader.ImpalaLoaderConfig.periodToLoad: “P13M” val minTime = currentDay.minus(config.feedPeriod)
listFiles.filter(file => file.eventDateTime isAfter minTime)
Solution loader.ImpalaLoaderConfig.periodToLoad: “P5D” loader.ImpalaLoaderConfig.periodToLoad: “P1M” loader.ImpalaLoaderConfig.periodToLoad: “P13M” …
Story #2. Missing data
Data ingestion Data from Partner X Data costs attribution Extractor
Problem XX Advertiser ID, Language, XX Device Type, …, XX
Media Cost (USD) X Advertiser ID, Language, X Device Type, …, X Media Cost (USD)
Solution • Rename old columns • Reload data for the
week
Solution val colRegex: Regex = “””X (.+)“””.r val oldNewColumnsMapping =
df.schema.collect { case oldColdName@colRegex(pattern) => (oldColName.name, (“XX “ + pattern)) } oldNewColumnsMapping.foldLeft(df) { case (data, (oldName, newName)) => data.withColumnRenamed(oldName, newName) }
XX Advertiser ID, Language, XX Device Type, …, XX Media
Cost (USD) Solution
Story #3. Divide and conquer
Problem processing_time part-*.parquet filtering aggregations created part-*.parquet
• Slow processing • Large parquet files • Failing job
that consumes lots of resources Problem
• Write new partitioned state • Run downstream jobs with
smaller states • Generate seed partition column - xxhash64(fullUrl, domain) Solution
processing_time part-*.parquet created bucket=0 part-*.parquet part-*.parquet … bucket=9 part-*.parquet part-*.parquet
processing_time part-*.parquet Solution
Story #4. Catch the evolution train
Data organisation evolution
Problem • Missing columns from the source • Impala to
Databricks migration speed • Dependency with another team • Unhappy users
Log-level data Mapper Ingestor Transformer Data costs calculator Data costs
attribution
Data costs attribution Data costs attribution Data extractor Impala loader
Data costs attribution Data extractor Impala loader Data costs attribution
Solution XX Advertiser ID, Language, XX Device Type, …, XX
Partner Currency, XX CPM Fee (USD) XX Advertiser ID, Language, XX Device Type, …, XX Media Cost (USD) 26 columns 82 columns
Solution Data extractor New ingestion job
//final step is writing the data df.write .partitionBy(“event_date”, “event_hour”) .mode(SaveMode.Overwrite)
.parquet(dstPath) Solution
Why this solution doesn’t work data_feed clicks.csv.gz views.csv.gz activity.csv.gz event_date
clicks1.parquet clicks2.parquet
Impressions Clicks Conversions Attribution data source
Solution impressions clicks conversions clicks.csv.gz views.csv.gz activity.csv.gz
Story #5. Cleanup time
Corrupted data Data from Partner X Ingestor
Corrupted data Data from Partner X Ingestor IllegalArgumentException: Can't convert
value to BinaryType data type
Solution • Adjust pipeline • Reload data for 3 days
on S3 • Relaunch Databricks autoloader
Current solution impressions videoevents conversions impressions conversions Clicks clicks videoevents
Current solution impressions conversions clicks videoevents
Better solution impressions videoevents conversions impressions conversions clicks clicks videoevents
Conclusions
2. Observability is the key 4. Plan major changes carefully
1. Set up clear expectations with stakeholders Prevention mechanisms 3. Distribute data transformation load
2. Errors can be prevented 4. Data evolution is hard
1. Data setup is always changing Conclusions 3. There are multiple approaches with different tools
None
dead_ fl owers22 roksolana-d roksolanadiachuk roksolanad My contact info