Upgrade to Pro
— share decks privately, control downloads, hide ads and more …
Speaker Deck
Features
Speaker Deck
PRO
Sign in
Sign up for free
Search
Search
Local Models for Coding
Search
Eugene Oskin
May 18, 2026
Programming
18
0
Share
Embed
Copy iframe code
Copy JS code
Copy link
Start on current slide
Local Models for Coding
Eugene Oskin
May 18, 2026
More Decks by Eugene Oskin
See All by Eugene Oskin
REST API. Django, Ruby on Rails, Play! Framework
evgeneoskin
0
100
Introduction to gRPC
evgeneoskin
0
110
GrailInventory – Advanced Backend Development
evgeneoskin
0
48
Bracing Calculator
evgeneoskin
1
80
emotional intelligence, part 2
evgeneoskin
0
49
Office temperature
evgeneoskin
0
45
Parse platform
evgeneoskin
0
110
Hubot
evgeneoskin
0
61
An introduction to iOS development
evgeneoskin
0
60
Other Decks in Programming
See All in Programming
PyConJP2026_wat_Python × Signal Processing: How to Draw Pictures with Sound Using Spectrogram Art
wat
0
680
Claude Codeを組織的に動かして月400PRを実現した話
happy_ryo
0
260
AIの中の人になってみる
htkym
0
170
The Good Stuff, Not the Slop: Engineering High-Quality Android Apps with Modern AI Tooling
danybony
1
200
GKE アップグレード前に知っておきたい Blue/Green と PDB の関係
stkk
0
160
AI に Inclusive UI を書かせよう — Design Rules Skill で Compose UI を作り直す
theoriatec2024
1
460
Swift愛好会と私(ウホーイ) / Swift Fan Club and Uhooi
uhooi
0
140
信頼性の目標を誰も求めてない
shubox
0
490
typoなんかねぇよ
raspython3
0
670
PHPプロジェクトの結合バランスを可視化する #php_night
kajitack
0
230
XHTMLが残したもの
yosuke_furukawa
PRO
2
440
自分的「カンファレンスの楽しみ方」
syumai
0
200
Featured
See All Featured
Impact Scores and Hybrid Strategies: The future of link building
tamaranovitovic
0
430
Mobile First: as difficult as doing things right
swwweet
225
10k
The AI Search Optimization Roadmap by Aleyda Solis
aleyda
1
6.2k
技術選定の審美眼(2025年版) / Understanding the Spiral of Technologies 2025 edition
twada
PRO
120
120k
The World Runs on Bad Software
bkeepers
PRO
72
12k
Test your architecture with Archunit
thirion
2
2.4k
Avoiding the “Bad Training, Faster” Trap in the Age of AI
tmiket
0
230
Agile that works and the tools we love
rasmusluckow
331
22k
Save Time (by Creating Custom Rails Generators)
garrettdimon
PRO
32
4.7k
JAMstack: Web Apps at Ludicrous Speed - All Things Open 2022
reverentgeek
1
590
The Limits of Empathy - UXLibs8
cassininazir
1
640
For a Future-Friendly Web
brad_frost
183
10k
Transcript
Local Models for Coding 1
Who am I • Local AI enthusiast running local models
on my MacBook • Building iowise.dev • ex-CTO of Termius SSH Client • 14 years working in software • Building backend, frontend, iOS, Android, and embedded
Running Models 3
OPENAI_API_BASE="https: / / chivalry - confront - untried.ngrok - free.dev:1234/v1"
OPENAI_API_KEY=lm - studio codex - - oss - - model qwopus3.6-35b - a3b - v1-mlx ANTHROPIC_BASE_URL="https: / / chivalry - confront - untried.ngrok - free.dev:1234/v1" ANTHROPIC_AUTH_TOKEN=lm - studio ANTHROPIC_API_KEY="" claude - - model qwopus3.6-35b - a3b - v1-mlx Connect Your Agent • Add model: qwopus3.6-35b-a3b-v1-mlx • Add OpenAI API key: lm-studio • Add OpenAI URL: https://chivalry-confront-untried.ngrok-free.dev:1234/v1 4
Terms • Model • Dense vs MoE • Agents •
Quantization • KV Cache • Inference engines • Prefill • Distillation • Speculations
-A3B -35B Model Names Qwen3.6 Model Family Total Number of
Parameters -Q4_K_S.gguf 6
Models • Qwen 3.x • Nemotron 3 • Gemma 4
• Fine-tuned
Where Models Live huggingface.co 🤗 8
Selecting a model • Add hardware to 🤗 • Browse
• Give it a go • Find eval • Iterate
How to run a model • LM Studio – https://lmstudio.ai/
• Ollama – https://ollama.com/ • Llama.cpp – the cutting edge of inference
Memory Math Total memory ≈ parameters × bytes/parameter + 2
× num_layers × num_kv_heads × head_dim × seq_len × batch_size × bytes/element + runtime overhead
Memory Math Total memory ≈ parameters × bytes/parameter + 2
× num_layers × num_kv_heads × head_dim × seq_len × batch_size × bytes/element + runtime overhead Start simple with LM Studio model loading Guardrails
-A3B -35B Qwen3.6 Model Family Total Number of Parameters -Q4_K_S.gguf
Active parameters 13
Dense VS Mixture of experts • Dense – “smarter” and
“slower” • MoE – “faster” but can be “dummier”
Faster Inference 15
-A3B -35B Qwen3.6 Model Family Total Number of Parameters -Q4_K_S.gguf
Active parameters Quantization File format 16
Quantization • F32 – original • F16 – 2x smaller
• Q8_0 – 4x smaller, similar quality • Q6_K • Q5_K_S, Q5_K_M, etc • Q4_K_S, Q4_K_M, etc – 8x smaller, the edge of quality • Q3_K_S, Q3_K_M, Q3_K_L, etc • Q2_K • 1-bit models
Quantization • ⬆ Inference • ⬇ RAM • ⬇ Disk
space • ⬇ KV Cache
Distillation • The original distillation requires full probability distribution of
output tokens • Modern distillation: training on synthetic data of a smart model
KV Cache It fixes a flaw in the LLM computation
model 20
KV Cache It fixes a flaw in the LLM computation
model 21
KV Cache It fixes a flaw in the LLM computation
model 22
KV Cache It fixes a flaw in the LLM computation
model 23
Inference stages 24
Inference stages 25
Speculations – Probabilistic speedup 26
Coding Agents • OpenCode • Pi • Goose • Hermes
Agent • OpenClaw • OpenAI Codex • Claude Code
Do you want to hear more? 28
Links • https://bbycroft.net/llm • https://hfviewer.com/Qwen/Qwen3.6-35B-A3B-FP8 • https://hfviewer.com/nvidia/Nemotron-3-Nano-Omni-30B-A3B-Reasoning- BF16 • http://hfviewer.com/nvidia/Nemotron-Cascade-2-30B-A3B
• https://huggingface.co/spaces/mlx-community/mlx-my-repo