Upgrade to Pro
— share decks privately, control downloads, hide ads and more …
Speaker Deck
Sign up for free
Menu
Search
Features
All features
Private URLs
Password Protection
Custom URLS
Scheduled publishing
Remove Branding
Restrict embedding
Deck Collections
Notes
Features
All features
Private URLs
Password Protection
Custom URLS
Scheduled publishing
Remove Branding
Restrict embedding
Deck Collections
Notes
Explore
Featured decks
Featured speakers
Programming
Technology
Storyboards
Explore
Featured decks
Featured speakers
Programming
Technology
Storyboards
Pricing
Search
Sign in
Sign up for free
ELYA-japanese-Llama-2-7bを Rust(WASM)で動かしてみた
Search
clouddev-code
January 30, 2024
830
1
Share
Embed
Copy iframe code
Copy JS code
Copy link
Start on current slide
ELYA-japanese-Llama-2-7bを Rust(WASM)で動かしてみた
clouddev-code
January 30, 2024
More Decks by clouddev-code
See All by clouddev-code
エージェント型プラットフォーム設計.pdf
cloudevcode
0
81
ECSでのinit_containerを利用したOpen_Telmetry自動計装.pdf
cloudevcode
0
180
microVMsのユースケースを考える.pdf
cloudevcode
0
57
cdk8s_Helm_どちらを選ぶべきか_rev.pdf
cloudevcode
0
250
Regional_NAT_Gatewayについて_basicとの違い_試した内容スケールアウト_インについて_IPv6_dual_networkでの使い分けなど.pdf
cloudevcode
1
940
Grafana_LokiをECS_Fargateで構築する観点公開版.pdf
cloudevcode
0
68
ADK_for_Java.pdf
cloudevcode
1
120
initContainerをECSで実現したい.pdf
cloudevcode
0
69
VPC_Lattice検討したが_採用しなかった話.pptx.pptx.pdf
cloudevcode
0
40
Featured
See All Featured
RailsConf & Balkan Ruby 2019: The Past, Present, and Future of Rails at GitHub
eileencodes
141
35k
WENDY [Excerpt]
tessaabrams
14
40k
Rebuilding a faster, lazier Slack
samanthasiow
85
9.7k
Jess Joyce - The Pitfalls of Following Frameworks
techseoconnect
PRO
1
420
Taking LLMs out of the black box: A practical guide to human-in-the-loop distillation
inesmontani
PRO
3
2.4k
[Rails World 2026] Durable orchestration on Rails: from continuation to workflow
palkan
1
400
Fight the Zombie Pattern Library - RWD Summit 2016
marcelosomers
234
18k
Refactoring Trust on Your Teams (GOTO; Chicago 2020)
rmw
35
3.8k
Google's AI Overviews - The New Search
badams
0
1.6k
Skip the Path - Find Your Career Trail
mkilby
1
240
Gemini Prompt Engineering: Practical Techniques for Tangible AI Outcomes
mfonobong
2
560
GraphQLとの向き合い方2022年版
quramy
50
15k
Transcript
ELYA-japanese-Llama-2-7bを Rust(WASM)で動かしてみた。 生成AI新年会2024@GMO Yours・フクラス 2024/01/30 #生成AI新年会 1
About us Soushi Hiruta 2 https://www.totalsolution.biz/ X(twitter) web_se Bluesky clouddevcode.bsky.social
github clouddev-code Zenn clouddevcode コンテナを中心にマイクロサービス基盤の構築、運用を行いつつ、 GenAIのキャッチアップを行っています。 Container、eBPF、GenAI
Zennにも検証したことをアップしています 3
Agenda ▸ LLMはGPUなどのComputeリソースを消費する ▸ WASM Runtimeのパフォーマンス ▹ Java等との違い ▹ 初期化プロセスとの違い
▸ WasmEdgeについて ▸ WASM-NN plugin ▸ ELYZA-japanese-Llhma-2-7b Model ▸ 上記モデルをRustで動かす点の注意点 ▸ デモ ▸ まとめ ▸ Q&A 4
5 LLMはComputeリソースを消費する https://xtech.nikkei.com/atcl/nxt/column/18/00989/091300127/
Python performance 6 There’s plenty of room at the Top:
What will drive computer performance after Moore’s law? https://www.science.org/doi/10.1126/science.aam9744
java, Python との違い 7 • Ahead-of-Time (AOT)は、実行前にバイトコードをマシンコードに変換 して最適化する • Java
は実行中にコンパイルされる。一度しか利用されないケースと かには向かない 対比されるものとしてDocker Engineはどうか
Docker Engineの初期化プロセス 8 • コントロールグループ (cgroup) • Rootfsのセットアップ これを終わらせたあとでないとアプリケーションを実行する ことができない
Performance advantages of WASM 9
初期化プロセス 10 • WASM Runtimeは主にアプリケーションバイナリを実行し、不要なファ イルシステム全体をマウントすることを回避する
11 WasmEdge Bring the cloud-native and serverless application paradigms to
Edge Computing • High performance • WASI-like Extensions • JavaScript Support • Cloud Native Management Orchestration • Cross-platform Support • Eas Extensibility • Easy to Embed into a Host Application
12 WasmEdge https://www.youtube.com/watch?v=BIgVM18UVIE
WASM-NN plugin 13 WasmEdge runtimes supports open-source LLMs through its
GGML plugin
ELYZA-japanese-Llama-2-7b 14 GPT-3.5 (text-davinci-003)に匹敵、日本語の公開モデル野中では最高 水準 約180億トークンの日本語テキストを追加 OSCARやWikipedia等に含まれる日本語テキストデータ
wasmedgeを動かすまでのポイント 15 https://github.com/second-state/LlamaEdge/blob/main/models.md • llama-api-server.wasmは最新のものを利用 ◦ 1/4にggmal pluginがリリースされている • メモリ8G程度だと、—ctx-size
オプション必須
デモ 16
まとめ 17 • WASMはDocker Engineと比較してもオーバーヘッドが少ない • GGMAL pluginは使って、OSS LLMなOpenAI ChatCompletion
互換なAPIを 構築できる • LlamaEdge 0.2.9が4h前にリリース(Phi-2などに対応)されるなど、アップ デートも活発です
Appendix 18 • WasmEdgeRuntime https://wasmedge.org/ • WasmEdge Provides a Better
Way to Run LLMs on the Edge https://www.secondstate.io/articles/wasmedge-ggml-plugin/ • WASM Runtimes vs. Containers: Cold Start Deplays (Part 1) https://levelup.gitconnected.com/wasm-runtimes-vs-containers-per formance-evaluation-part-1-454cada7da0b • Metaの「Llama 2」をベースとした商用利用な日本語LLMを公開しまし た。 https://note.com/elyza/n/na405acaca130 • GGUF Models • https://github.com/second-state/LlamaEdge/blob/main/models.md • ELYZA-japanese-Llama-2-7bをM1 Mac上でRustで動かす