Upgrade to Pro — share decks privately, control downloads, hide ads and more …

ローカル大規模言語モデル いまどこまで動く? 2026年6月版

Avatar for ayane ayane
July 08, 2026

ローカル大規模言語モデル いまどこまで動く? 2026年6月版

2026年6月18日のGenAIge Cafe #8「 https://genaige.connpass.com/event/396761/ 」で話した資料です

Avatar for ayane

ayane

July 08, 2026

More Decks by ayane

Other Decks in Technology

Transcript

  1. 2026年6⽉時点の、現実的な3つの層 モデルの規模 (例え:本の厚さ) ⽬安 必要なビデオメモリ感 初⼼者向けの理解 7B 〜 8B 前後

    ⼩さめ 6GB〜8GBから動き始める (12GB〜16GBあると安⼼) まず試すならここ! 24B 〜 32B 前後 中くらい 24GB〜32GBが実⽤ライン かなり実⽤的で⾯⽩い 70B 以上 ⼤きい 48GB〜80GB以上など (または複数台の連携が必要) 個⼈の趣味を超え、 ワークステーションやサーバー寄り
  2. ⽇本語でゆるく試すなら、この3つ! Qwen 3 系 公式情報で、119の⾔語と⽅⾔を サポート。多⾔語に強いモデルと して世界的に⼈気です。 Gemma 4 系

    公式情報で140以上の⾔語をサ ポート。様々なサイズがあり、環 境に合わせて選びやすいです。 Mistral Small 3.2 公式のモデル説明で「⽇本語対 応」を明記。中くらいのサイズ で、バランスの良さが光ります。
  3. どんなパソコンなら楽しめる?(ざっくり⽬安) ※情報を少し圧縮して軽くする「4ビット量⼦化」を⾏った場合の⽬安です。利⽤ソフト等で快適さは変わります。 1 まず試す ⼩さめのモデル(7B〜 8B)を⼿元で動かして みる⼊⾨編。 2 ちゃんと楽しむ ⼩さめモデルは快適。

    中くらい(24B前後) も⼯夫次第で試せま す。 3 本格的に遊ぶ Gemma 4 26B A4B QATなど中くらいが視 野に。例: RTX 4090, RTX 5090等 4 別世界‧⼤型 70B以上や、複雑な専 ⾨家連携モデル向け。 例: Radeon PRO W7900等
  4. ⼀次情報URL⼀覧 モデル公式‧モデルカード 実⾏環境‧量⼦化 / ハードウェア仕様 Meta Llama 4 公式ブログ https://ai.meta.com/blog/llama-4-multimodal-intelligence/

    Llama 3.1 8B Instruct model card https://huggingface.co/meta-llama/Llama-3.1-8B-Instruct Llama 3.3 70B Instruct model card https://huggingface.co/meta-llama/Llama-3.3-70B-Instruct Mistral Small 3.2 24B https://huggingface.co/mistralai/Mistral-Small-3.2-24B-Instruct-2506 Gemma 4 26B A4B https://huggingface.co/google/gemma-4-26B-A4B Gemma 4 26B A4B QAT https://lmstudio.ai/models/google/gemma-4-26b-a4b-qat Qwen 3 公式ブログ https://qwenlm.github.io/blog/qwen3/ Qwen3.6 27B https://huggingface.co/Qwen/Qwen3.6-27B Qwen3 30B-A3B model card https://huggingface.co/Qwen/Qwen3-30B-A3B llama.cpp 公式GitHub https://github.com/ggml-org/llama.cpp vLLM hardware support https://docs.vllm.ai/en/v0.9.2/features/quantization/supported_hardware.html Transformers bitsandbytes docs https://huggingface.co/docs/transformers/quantization/bitsandbytes PyTorch Gemma 3 27B INT4 benchmark https://huggingface.co/pytorch/gemma-3-27b-it-INT4 NVIDIA Llama 3.1 8B ONNX INT4 page https://huggingface.co/nvidia/Meta-Llama-3.1-8B-Instruct-ONNX-INT4 Ollama MLX performance https://ollama.com/blog/mlx-performance NVIDIA GeForce RTX 4090 specs https://www.nvidia.com/en-us/geforce/graphics-cards/40-series/rtx-4090/ AMD Radeon PRO W7900 specs https://www.amd.com/en/products/graphics/workstations/radeon-pro/w7900.html NVIDIA ⽐較 GeForce グラフィックス カード https://www.nvidia.com/ja-jp/geforce/graphics-cards/compare/