Upgrade to Pro
— share decks privately, control downloads, hide ads and more …
Speaker Deck
Sign up for free
Menu
Search
Features
All features
Private URLs
Password Protection
Custom URLS
Scheduled publishing
Remove Branding
Restrict embedding
Deck Collections
Notes
Features
All features
Private URLs
Password Protection
Custom URLS
Scheduled publishing
Remove Branding
Restrict embedding
Deck Collections
Notes
Explore
Featured decks
Featured speakers
Programming
Technology
Storyboards
Explore
Featured decks
Featured speakers
Programming
Technology
Storyboards
Pricing
Search
Sign in
Sign up for free
CRubyプロダクトにおけるembulkの活用法
Search
Tomohiro Hashidate
May 16, 2017
Programming
780
1
Share
Embed
Copy iframe code
Copy JS code
Copy link
Start on current slide
CRubyプロダクトにおけるembulkの活用法
embulk meetup #03
Tomohiro Hashidate
May 16, 2017
More Decks by Tomohiro Hashidate
See All by Tomohiro Hashidate
Ruby::Boxでできること、Refinementsでできること
joker1007
3
520
Do Ruby::Box dream of Modular Monolith?
joker1007
1
1.7k
ReproでのicebergのStreaming Writeの検証と実運用にむけた取り組み
joker1007
0
820
マイクロサービスへの5年間 ぶっちゃけ何をしてどうなったか
joker1007
23
11k
Quarkusで作るInteractive Stream Application
joker1007
0
330
今改めてServiceクラスについて考える 〜あるRails開発者の10年〜
joker1007
25
29k
rubygem開発で鍛える設計力
joker1007
5
1.5k
実践Kafka Streams 〜イベント駆動型アーキテクチャを添えて〜
joker1007
3
1.5k
本番のトラフィック量でHudiを検証して見えてきた課題
joker1007
2
1.4k
Other Decks in Programming
See All in Programming
Vue Fes Japan 2026 タイムテーブル徹底解説
448jp
1
590
JRuby: Past, Present, and Future
headius
0
220
[ハンズオン]AIへの指示だけで「五目並べ」を作ってみよう
satoshi256kbyte
1
320
SREの越境 / SRE Collaboration
y0hgi
2
290
Agents on Rails - Rails at Scale 2026
irinanazarova
0
310
Agentic Software Factoryに、すごく賢いIF文を。 / super smart IF statement into the Agentic Software Factory.
rkaga
1
460
KiroのSpecで「五目並べ」を作ってみる
satoshi256kbyte
1
330
GemmaをJevのように使ってみる / Use Gemma like Jev
kishida
5
520
FreeBSDでZabbixを動かす
kenkino
0
350
更なる可用性を求めて、5年間運用したKotlinのアプリケーションをGoでリプレイスする話
ken_tunc
0
430
速習iPhone Duo対応
yuukiw00w
2
930
Go × SIMDで高速化するベクトル検索 ~ルーフラインモデルでSIMDが効く境界を探れ! ~
po3rin
1
6k
Featured
See All Featured
The Limits of Empathy - UXLibs8
cassininazir
1
700
Designing Powerful Visuals for Engaging Learning
tmiket
1
590
Design and Strategy: How to Deal with People Who Don’t "Get" Design
morganepeng
133
20k
Git: the NoSQL Database
bkeepers
PRO
432
67k
How to Get Subject Matter Experts Bought In and Actively Contributing to SEO & PR Initiatives.
livdayseo
0
200
Typedesign – Prime Four
hannesfritz
42
3.2k
Side Projects
sachag
456
43k
The untapped power of vector embeddings
frankvandijk
2
1.9k
Facilitating Awesome Meetings
lara
57
7.2k
What’s in a name? Adding method to the madness
productmarketing
PRO
24
4.2k
Jess Joyce - The Pitfalls of Following Frameworks
techseoconnect
PRO
1
430
Build your cross-platform service in a week with App Engine
jlugia
234
19k
Transcript
CRuby プロダクトにおける embulk の活用法 @joker1007
self.inspect @joker1007 Repro inc. CTO ( 要は色々やる人) Ruby/Rails uentd/embulk ←
今日はこの辺 Docker/ECS Bigquery/EMR/Hive/Presto
Repro のサービス モバイルアプリケーションの行動トラッキング 分析結果の提供と、それと連動したマーケティン グの提供 大体Ruby ・Rails でほぼAWS 上で稼動している Docker
やterraform 等も活用している 会社規模の割にデータ量が多い。 そのためデータエンジニアリングも必要。
開発・メンテしてるもの embulk embulk- lter-ruby_proc embulk-output-in uxdb embulk-parser(formatter)-avro embulk-output-s3_per_record その他 uent-plugin-bigquery
rukawa ( 自作のワークフローエンジン)
embulk の主な用途 Bigquery で集計した結果の整形、取得 ( 日次バッ チ) データの洗い替え ( データメンテナンス)
embulk の採用理由 CRuby より高速に処理できる かつRuby でプラグインを書くことでアドホックな 加工ができる プラグイン管理にRubyist にとって馴染みのある Bundler
が使える Gem le.lock によるバージョンロックが楽 プラグインインストールの方法がプロダクトと 揃う JRuby やプラグイン機構がRubyist フレンドリーな点 が良い。
バッチ処理に組込むために必要なもの 自動実行のためのスケジューラ -> Rundeck 実行日付を元に設定をパラメーター化 -> yaml_master 処理同士の依存関係や並列数を定義できるワーク フローエンジン ->
rukawa digdag があればいい。 が、現在、digdag は使っていないw digdag リリースの前に設計・構築したので。
rukawa については以下を参照 http://joker1007.github.io/slides/introduce_rukawa/slid es/index.html
embulk 利用例詳細 1. 更新が発生するデータをembulk でBq に転送 2. uentd で蓄積しているログと結合しBq で集計
3. 集計後のデータとそれに紐付くユーザーID をAVRO でexport 4. GCS からS3 にembulk でデータを転送する 5. S3 に転送したAVRO ファイルをembulk でRDB に import 6. S3 に転送したAVRO ファイルをembulk でレコード 単位にばらして別バケットに保存
Bq へのデータ転送 日次バッチで必要なデータを都度丸ごと上書き 更新のある小規模のデータで羃等性を担保するた め embulk-output-bigquery を普通に利用している
Bq でのデータ集計 いくつかのSQL ジョブをワークフローエンジンで定 義して実行 処理が終わった側からexport してはembulk でデー タを転送する
GCS からS3 へのデータ転送 embulk-input-gcs とembulk-output-s3 を使用 ファイルフォーマットにAVRO を利用するため以下 のプラグインを開発 embulk-parser-avro
embulk-formatter-avro
embulk-parser-avro in: type: file path_prefix: "items" parser: type: avro avsc
: "./item.avsc" columns: - {name: "id", type: "long"} - {name: "name", type: "string"} - {name: "flag", type: "boolean"} - {name: "price", type: "long"} - {name: "item_type", type: "string"} - {name: "tags", type: "json"} - {name: "options", type: "json"} out: type: stdout
AVRO スキーマ サンプル { "type" : "record", "name" : "Item",
"namespace" : "example.avro", "fields" : [ {"name": "id", "type": "int"}, {"name": "name", "type": "string"}, {"name": "flag", "type": "boolean"}, {"name": "spec", "type": { "type": "record", "name": "item_spec", "fields" : [ {"name" : "key", "type" : "string"}, {"name" : "value", "type" : ["string", "null"]} ]} } ] }
AVRO を使う理由 Bq からexport した時のデータ型の問題 JSON だとINT が文字列になる INT の配列も文字列になる
スキーマを後から変えられる スキーマ側からデフォルト値を差し込める Hadoop エコシステムで処理しやすい cf. http://avro.apache.org/docs/current/spec.html
レコードをばらしてS3 に書き込む サービス上の要請による embulk-output-s3_per_record を利用 Page をadd する時に都度S3 に書き込む 効率がめっちゃ悪いので遅いが仕方なく
embulk-output-s3_per_record con g out: type: s3_per_record bucket: your-bucket-name key: "sample/${id}.txt"
mode: multi_column serializer: json data_columns: [id, payload] output {"id": 5, "payload": "foo"}
設定ファイルの生成とembulk の実行 yaml_master というgem を作り設定ファイルを生成 ERB を利用してRuby からプロパティを突っ込ん でyaml を出力する
Liquid より表現力が高い CRuby のプロダクトが持つ情報を元に直接生成 しやすい CRuby からの呼び出しはpopen3 を利用してプロセ スを起動する Rukawa で呼び出しのためのヘルパーを書いてい る
Embulk を効果的に使うには システムに組込むには、いくつか周辺のヘルパー を用意する必要がある 自分達が使うフォーマットに合わせてプラグイン を作れる様にしておく パフォーマンス要求がきつくないデータ加工には embulk- lter-ruby_proc をオススメする
プラグイン開発で意識していること できるだけ機能を限定する カラムの操作など lter でやれることは単体の lter に任せる データのparse, encode はJava
の方が良い ネストしたデータ構造を扱うのは面倒だが マルチスレッドで動作することを意識する 特にRuby プラグイン 実行フェイズ毎にシリアライズを挟む場合がある ことを知っておく