Upgrade to Pro
— share decks privately, control downloads, hide ads and more …
Speaker Deck
Sign up for free
Menu
Search
Features
All features
Private URLs
Password Protection
Custom URLS
Scheduled publishing
Remove Branding
Restrict embedding
Deck Collections
Notes
Features
All features
Private URLs
Password Protection
Custom URLS
Scheduled publishing
Remove Branding
Restrict embedding
Deck Collections
Notes
Explore
Featured decks
Featured speakers
Programming
Technology
Storyboards
Explore
Featured decks
Featured speakers
Programming
Technology
Storyboards
Pricing
Search
Sign in
Sign up for free
メモリウォールを超えて:キャッシュメモリ技術の進歩
Search
kawayu
April 13, 2025
Programming
3.7k
0
Share
Embed
Copy iframe code
Copy JS code
Copy link
Start on current slide
メモリウォールを超えて:キャッシュメモリ技術の進歩
kawayu
April 13, 2025
More Decks by kawayu
See All by kawayu
AIエージェントの隔離技術の徹底比較
kawayu
0
580
Other Decks in Programming
See All in Programming
App Storeの外へ──日本のiOSサイドローディング入門 for iOSDC Japan 2026
yuukiw00w
0
200
What We Talk About When We Talk About XP
m_seki
2
620
そのリトライ、死んだコネクションを使い回していませんか ── GoのHTTPクライアントとHTTP/2を実プロダクト障害から学び直す
myus4a
0
130
Are APIs Still Relevant in the AI Era?
soyuka
0
170
The Past, Present, and Future of Enterprise Java
ivargrimstad
0
350
標準パッケージに uuid が追加された 背景から見る Go らしい意思決定 / go_127_uuid_decision
convto
5
7k
スマート反転とウェブアクセシビリティ
camiha
0
210
FreeBSDでZabbixを動かす.pdf
kenkino
0
250
マイコン向けの軽量Ruby「PicoRuby」で各種デバイスを制御するネイティブアプリの実現手法
bash0c7
0
390
[DroidKaigi 2026] Bring your own phones to Gradle Managed Devices
f2lk
0
120
高専キャリア LT 発表内容
crysta1221
6
5.7k
AI × TiDD / 2026.09.05 Redmine 大阪
tokudiro
1
160
Featured
See All Featured
Balancing Empowerment & Direction
lara
6
1.3k
Rebuilding a faster, lazier Slack
samanthasiow
85
9.6k
The #1 spot is gone: here's how to win anyway
tamaranovitovic
3
1.2k
Conquering PDFs: document understanding beyond plain text
inesmontani
PRO
4
3.1k
HU Berlin: Industrial-Strength Natural Language Processing with spaCy and Prodigy
inesmontani
PRO
0
700
Odyssey Design
rkendrick25
PRO
2
810
How Fast Is Fast Enough? [PerfNow 2025]
tammyeverts
3
880
Breaking role norms: Why Content Design is so much more than writing copy - Taylor Woolridge
uxyall
1
400
Groundhog Day: Seeking Process in Gaming for Health
codingconduct
0
350
16th Malabo Montpellier Forum Presentation
akademiya2063
PRO
0
380
We Analyzed 250 Million AI Search Results: Here's What I Found
joshbly
1
1.9k
Templates, Plugins, & Blocks: Oh My! Creating the theme that thinks of everything
marktimemedia
31
2.9k
Transcript
メモリウォールを超えて: キャッシュメモリ技術の進歩 2025/4/12 開発合宿@鳴子温泉 @kawayu_u
メモリウォール • メモリっておそいな〜って思ったことありますよね?
メモリウォール • メモリっておそいな〜って思ったことありますよね? • CPUの性能向上が顕著になった1990年代にメモリウォール問題 が提起 • CPUの性能向上に対しメモリ性能の向上は鈍い
メモリとCPUの性能推移 Source: Carlos Carvalho ,The Gap between Processor and Memory
Speeds
メモリとCPUの性能推移 Source: Carlos Carvalho ,The Gap between Processor and Memory
Speeds https://www.cs.columbia.edu/~martha/ courses/3827/au14/advanced- topics.pdf
メモリとCPUの性能推移 メモリウォールは改善するどころか性能差が広がりより顕著に
CPUキャッシュメモリの登場 • CPUの近くに配置→電気信号の伝搬時間が短い • SRAMを使用→高速アクセスが可能 • SRAM(Static Random Access Memory)とは
• DRAMと比べ読み書きが高速 CPU メモリ (DRAM) キャッシュ メモリ (SRAM)
CPUキャッシュメモリ https://www.fujitsu.com/jp/products/computing/servers/unix/term/cache/index.html
CPUキャッシュの歴史 CPU名 発売年 世代/シリーズ L1キャッシュ L2キャッシュ L3キャッシュ Intel 4004 1971年
なし なし なし Intel 8008 1972年 なし なし なし Intel 8080 1974年 なし なし なし Intel 8086 1978年 x86-16 なし (6Bプリフェッチ) なし なし Intel 80286 1982年 x86-16 なし なし なし Intel 80386 1985年 x86-32 なし (16Bプリフェッチ) なし なし Intel 80486 DX 1989年 x86-32 8 KB なし なし Intel Pentium 1993年 Pentium 8 KB I + 8 KB D 最大512 KB (オンボード) なし Intel Pentium Pro 1995年 Pentium Pro 8 KB I + 8 KB D 256 KB - 1 MB (オンパッケージ) なし Intel Pentium II 1997年 Pentium II 16 KB I + 16 KB D 512 KB (オフダイ) なし Intel Pentium III 1999年 Pentium III 16 KB I + 16 KB D 128 KB - 512 KB (オンダイ/オフダイ) なし Intel Pentium 4 2000年 NetBurst 8 KB D (トレースキャッシュI) 最大2 MB 最大2 MB (一部モデル)
現代のCPU https://jp.fujitsu.com/family/familyroom/syuppan/family/webs/serial-comp2/index.html 微細化によって演算ユニットが小さくなる キャッシュメモリに使用できる面積増加 キャッシュメモリの容量を増加 現在はキャッシュメモリに大きなスペースを使用
CPUキャッシュメモリ CPUってマルチコアですよね? L1, L2キャッシュメモリって共有できるんですか? https://www.fujitsu.com/jp/products/computing/servers/unix/term/cache/index.html
L3Cacheの登場 https://www.cs.columbia.edu/~martha/courses/3827/au14/advanced-topics.pdf 複数コアで共有するL3Cacheが登場
CPUキャッシュの歴史 CPU名 発売年 世代/シリーズ L1キャッシュ L2キャッシュ L3キャッシュ Intel Core 2
Duo E8400 2008年 Core 2 32 KB I + 32 KB D/コア 6 MB (共有) なし Intel Core i7-920 2008年 Nehalem (Core i) 32 KB I + 32 KB D/コア 256 KB/コア 8 MB (共有) Intel Core i7-2600K 2011年 Sandy Bridge (Core i) 32 KB I + 32 KB D/コア 256 KB/コア 8 MB (共有) Intel Core i7-3770K 2012年 Ivy Bridge (Core i) 32 KB I + 32 KB D/コア 256 KB/コア 8 MB (共有) Intel Core i7-4770K 2013年 Haswell (Core i) 32 KB I + 32 KB D/コア 256 KB/コア 8 MB (共有) Intel Core i7-5775C 2014年 Broadwell (Core i) 32 KB I + 32 KB D/コア 256 KB/コア 6 MB (共有) + 128 MB L4 Intel Core i7-6700K 2015年 Skylake (Core i) 32 KB I + 32 KB D/コア 256 KB/コア 8 MB (共有) Intel Core i7-7700K 2017年 Kaby Lake (Core i) 32 KB I + 32 KB D/コア 256 KB/コア 8 MB (共有) Intel Core i7-8700K 2017年 Coffee Lake (Core i) 32 KB I + 32 KB D/コア 256 KB/コア 12 MB (共有) Intel Core i7-9700K 2018年 Coffee Lake Refresh (Core i) 32 KB I + 32 KB D/コア 256 KB/コア 12 MB (共有) Intel Core i7-10700K 2020年 Comet Lake (Core i) 32 KB I + 32 KB D/コア 256 KB/コア 16 MB (共有) Intel Core i7-1065G7 2019年 Ice Lake (Core i) 32 KB I + 48 KB D/コア 512 KB/コア 8 MB (共有) Intel Core i7-11700K 2021年 Rocket Lake (Core i) 32 KB I + 32 KB D/コア 512 KB/コア 16 MB (共有) Intel Core i7-11850H 2020年 Tiger Lake (Core i) 32 KB I + 48 KB D/コア 1.25 MB/コア 24 MB (共有) Intel Core i7-12700K 2021年 Alder Lake (Core i) 80 KB (P)/64 KB (E) 2 MB (P)/4 MB (Eクラスタ) 25 MB (共有) Intel Core i7-13700K 2022年 Raptor Lake (Core i) 80 KB (P)/96 KB (E) 2 MB (P)/4 MB (Eクラスタ) 30 MB (共有)
CPUキャッシュの歴史 CPU名 発売年 世代/シリーズ L1キャッシュ L2キャッシュ L3キャッシュ AMD Athlon 64
X2 2005年 K8 64KB I + 64KB D/コア 512KB/コア - AMD Ryzen 7 1800X 2017年 Zen 64KB I + 32KB D/コア 512KB/コア 16MB (共有) AMD Ryzen 7 3800XT 2019年 Zen 2 64KB I + 32KB D/コア 512KB/コア 32MB (共有) AMD Ryzen 7 5800X 2020年 Zen 3 64KB I + 32KB D/コア 512KB/コア 32MB (共有) AMD Ryzen 7 5700X 2022年 Zen 3 64KB I + 32KB D/コア 512KB/コア 32MB (共有) AMD Ryzen 7 5700 2022年 Zen 3 64KB I + 32KB D/コア 512KB/コア 16MB (共有) AMD Ryzen 7 7800X3D 2023年 Zen 4 64KB I + 32KB D/コア 1MB/コア 96MB (共有)
キャッシュメモリを大きく • キャッシュメモリを大きくすればチップ性能は向上するが… • 大きなダイサイズ→コストUP • 動かなかったとき→歩留まり率低下 • 演算ユニットの集積化は頭打ちが近い
現代のCPU Apple Silicon(M1) • SoC(System On a Chip)を採用 • メモリはCPUとGPUで共用
• CPUとメモリの距離が近い • M1ではL3 Cacheは非搭載 • L2キャッシュとユニファイドメモリが共用 https://www.production-expert.com/production-expert-1/why-are-the-apple-m1-m1-pro-and-m1-max-chips-so-fast
現代のCPU AMD 3D V-Cache • AMDはチップレット(CCD) 戦略を採用 • CCDの垂直方向にL3 Cacheを搭載
• 大容量L3 Cacheにより性能を大幅向上 https://fuse.wikichip.org/news/5531/amd-3d-stacks-sram-bumplessly/
現代のCPU AMD 3D V-Cache CCD(Core Complex, and Die) https://www.amd.com/content/dam/amd/en/documents/epyc-business-docs/white-papers/221704010-B_en_4th-Gen-AMD-EPYC-Processor-Architecture---White-Paper_pdf.pdf
現代のCPU AMD 3D V-Cache CCD(Core Complex, and Die)のメリット https://www.amd.com/ja/technologies/zen-core.html
現代のCPU AMD 3D V-Cache https://fuse.wikichip.org/news/5531/amd-3d-stacks-sram-bumplessly/
実際に計測
実際に計測
実際に計測
実際に計測 0 0.0005 0.001 0.0015 0.002 0.0025 0.003 0 200
400 600 800 1000 1200 Sequential Access Random Access KB
実際に計測 MB 0 0.05 0.1 0.15 0.2 0.25 0.3 0
5 10 15 20 25 30 35 Sequential Access Random Access
計測結果と考察 • 12MBくらいからSequential と Randomに差が生じるように • 詳細は公表されていないがM1チップのL2 Cacheは12MBのよ うでそれを示す結果 •
https://en.wikipedia.org/wiki/Apple_M1 • 処理を予測してメモリからキャッシュメモリに事前にデータを コピーするプリフェッチ機能があるが12MB以下では予測でき ない場合でもキャッシュメモリにデータがあるため、処理時間 に差が生じていないと考えられる