DeepSeek-Coder-V2-Lite Instruct
DeepSeek
A 16B MoE code model with only 2.4B active per token — fast enough for inline completion while holding 128k of context.
DeepSeek-Coder-V2-Lite Instruct does not fit on CPU only — 32 GB DDR5.
CPU-only inference. Expect a few tokens per second at best.
- Very fast generation
- 128k context
- Supports 300+ languages
- Needs the full 16B in memory
- Custom licence
deepseek-coder-v2:16bRunning DeepSeek-Coder-V2-Lite Instruct on CPU only — 32 GB DDR5
| Quantization | Quality | Weights | KV cache | Total | ~tok/s | Fit |
|---|---|---|---|---|---|---|
| F16 | lossless | 29 GB | 1.7 GB | 32 GB | 13 | Won't run |
| Q8_0 | near-lossless | 16 GB | 1.7 GB | 18 GB | 25 | Not recommended |
| Q6_K | near-lossless | 12 GB | 1.7 GB | 14 GB | 33 | Not recommended |
| Q5_K_M | high | 10 GB | 1.7 GB | 13 GB | 38 | Not recommended |
| Q4_K_M | balanced | 8.8 GB | 1.7 GB | 11 GB | 44 | Not recommended |
| Q3_K_M | degraded | 7.2 GB | 1.7 GB | 9.6 GB | 55 | Not recommended |
KV cache is sized at 8,192 tokens. Longer contexts cost proportionally more — the recommendation above reserves room for a working context.
Local fine-tuning on this machine
- Base weights
- 8.2 GB
- Optimizer
- 356 MB
- Activations
- 500 MB
- Peak
- 11 GB
- Local fine-tuning needs a GPU or Apple Silicon. CPU training is not practical.
- Base weights
- 29 GB
- Optimizer
- 356 MB
- Activations
- 500 MB
- Peak
- 32 GB
- Local fine-tuning needs a GPU or Apple Silicon. CPU training is not practical.
Where this model runs
VRAM 32 GB · Q8_0 · 18 GB
VRAM 24 GB · Q5_K_M · 13 GB
VRAM 16 GB · Q4_K_M · 11 GB
VRAM 16 GB · Q4_K_M · 11 GB
VRAM 24 GB · Q5_K_M · 13 GB
VRAM 12 GB · Q4_K_M · 11 GB
VRAM 12 GB · Q4_K_M · 11 GB
VRAM 16 GB · Q4_K_M · 11 GB
VRAM 48 GB · F16 · 32 GB
VRAM 80 GB · F16 · 32 GB
Unified 128 GB · F16 · 32 GB
Unified 48 GB · Q8_0 · 18 GB
Unified 24 GB · Q4_K_M · 11 GB
Unified 192 GB · F16 · 32 GB
Unified 16 GB · Q4_K_M · 11 GB
VRAM 24 GB · Q5_K_M · 13 GB
VRAM 16 GB · Q4_K_M · 11 GB
VRAM 0 MB
VRAM 0 MB
Qwen2.5 14B Instruct
14.8B · Apache 2.0
The sweet spot for 24 GB cards. Meaningfully stronger reasoning than 7B while still fine-tunable locally with QLoRA.
Qwen3 14B
14.8B · Apache 2.0
The Qwen3 mid-size. Long context and a reasoning mode inside a footprint a 24 GB card handles comfortably.
Mistral Nemo 12B Instruct
12.2B · Apache 2.0
A 12B with a 128k context and the Tekken tokenizer, which compresses non-English text far better than Llama’s.
Gemma 3 12B Instruct
12.2B · Gemma Terms of Use
Multimodal and long-context in a 12B footprint. Interleaved local/global attention keeps the KV cache affordable.
Catalogue figures come from each model’s published card. ModelLM has not independently measured them.