Upgrade to Pro
— share decks privately, control downloads, hide ads and more …
Speaker Deck
Sign up for free
Menu
Search
Features
All features
Private URLs
Password Protection
Custom URLS
Scheduled publishing
Remove Branding
Restrict embedding
Deck Collections
Notes
Features
All features
Private URLs
Password Protection
Custom URLS
Scheduled publishing
Remove Branding
Restrict embedding
Deck Collections
Notes
Explore
Featured decks
Featured speakers
Programming
Technology
Storyboards
Explore
Featured decks
Featured speakers
Programming
Technology
Storyboards
Pricing
Search
Sign in
Sign up for free
AWSで始めるサーバーレスなデータ分析基盤
Search
afooooil
October 22, 2025
820
0
Share
Embed
Copy iframe code
Copy JS code
Copy link
Start on current slide
AWSで始めるサーバーレスなデータ分析基盤
JAWS-UG東京 ランチタイムLT会 #28(
https://jawsug.connpass.com/event/367465/
) で発表させていただいた資料です。
afooooil
October 22, 2025
More Decks by afooooil
See All by afooooil
workmuxで始めるClaude Codeの並列開発
afooooil
0
48
DynamoDBからS3(Icebergテーブル)へのZeroETLを行う
afooooil
1
120
退屈なことはAI_Agentにやらせよう
afooooil
0
250
Amazon Qとのより良い付き合い方を考える
afooooil
0
300
ZeroETLで始めるDynamoDBとS3の連携
afooooil
0
330
Featured
See All Featured
The innovator’s Mindset - Leading Through an Era of Exponential Change - McGill University 2025
jdejongh
PRO
1
350
Organizational Design Perspectives: An Ontology of Organizational Design Elements
kimpetersen
PRO
1
840
Designing for Timeless Needs
cassininazir
1
510
B2B Lead Gen: Tactics, Traps & Triumph
marketingsoph
0
250
Building Flexible Design Systems
yeseniaperezcruz
330
41k
Distributed Sagas: A Protocol for Coordinating Microservices
caitiem20
333
23k
Intergalactic Javascript Robots from Outer Space
tanoku
273
27k
Joys of Absence: A Defence of Solitary Play
codingconduct
1
530
Visualizing Your Data: Incorporating Mongo into Loggly Infrastructure
mongodb
50
10k
The World Runs on Bad Software
bkeepers
PRO
72
12k
The Illustrated Children's Guide to Kubernetes
chrisshort
51
53k
Lessons Learnt from Crawling 1000+ Websites
charlesmeaden
PRO
1
1.6k
Transcript
AWSではじめるサーバーレスな データ分析基盤 株式会社モリサワ 岡田 晃 JAWS-UG東京 ランチタイムLT会 #28
自己紹介 岡田 晃 / @afooooil 所属: 株式会社モリサワ ポジション: データエンジニア /
データサイエンティスト 最近興味のある技術: Apache Iceberg, DuckDB
データを分析、活用するための基盤。 • プロダクトなどからデータを収集して、 • 扱いやすい形に加工を行い、 • BIツールなどに連携し、活用する データ分析基盤とは? 収集 加工
活用
構築、運用にかかるコストを下げたかった。 • 自分のロールはデータ分析、活用 + 基盤の整備。 ◦ サーバーレスにすることで浮くリソースを分析業務に配分できる。 サーバーレスに絶対のこだわりがあるわけではなく、要件で必要になることがあれ ば、ECSやRedshiftを利用する。 •
コスト軽減が目的でありサーバーレス化はあくまで手段である。 なぜサーバーレス?
データ分析基盤のアーキテクチャ 収集: DynamoDBのPITRをLambdaでRawデータのS3へコピー。 加工: Athenaを用いて加工して、データレイクのS3へ移動。 データレイクではApache Icebergを利用。 活用: QuickSight(BIツール)をもちいてユーザーへデータ提供。
Apache Icebergとは? Apache IcebergとはOpen Table Formatのひとつ。 - 個々のファイルの集合をあたかも一つのテーブルのように扱える。 - 従来のデータレイクにある課題を解決する次世代のフォーマットとして注目され
ている。 嬉しい特徴の一つとして、レコードの追加、更新、削除を容易に効率的に行う ことがあげられる。 ここでは紹介しませんが、Icebergには他にも様々な魅力的な機能があり ます。
Icebergは何が嬉しいか? SQLを用いてS3上のデータの追加、更新、削除が行える • INSERT, MERGE, DELETEが使える • データソースの変更の差分を継続的にデータレイクに取り込むことも可能 • 一方でS3にあるファイルを直接触らなくて良い
そのためStepFunctionsでAthenaのクエリを定期的に実行するだけで データ変換が可能になる。 Redshiftの導入も視野に入れていたが、Icebergをデータレイクに導入した。 DynamoDBからSageMaker Lakehouse(Icebergテーブル)へのZeroETL も可能になっており導入に向けて検証中
まとめ • Lambda, Athena, QuickSightなどを使うことで、AWS上でサーバレスな データ分析基盤を構築することが可能。 • Apache Icebergではデータの追加、更新、削除が効率的、容易に行うことが できる。
• Apache Icebergをデータレイクに採用することでデータレイクの構築、運用 にかかるコストを低減できる。