Best local AI models for CPU only — 32 GB DDR5
No GPU. Small models are usable for batch work; anything interactive will feel slow. Retrieval still works well.
CPU fine-tuning is not practical at any useful model size.
Specification
VRAM32 GB
Usable for models22 GB
System RAM32 GB
Bandwidth83 GB/s
Local trainingNot practical
Operating systemswindows, linux, macos
Runs well
4 models fit this machine
Estimated
| Model | Params | Quantization | Memory | ~tok/s | Fit | Fine-tune |
|---|---|---|---|---|---|---|
| Phi-3.5 Mini Instruct | 3.8B | Q4_K_M | 5.7 GB | 28 | Runs with CPU offload | No |
| Llama 3.2 3B Instruct | 3.2B | Q4_K_M | 3.2 GB | 33 | Runs with CPU offload | No |
| SmolLM2 1.7B Instruct | 1.7B | Q4_K_M | 2.9 GB | 63 | Runs with CPU offload | No |
| Llama 3.2 1B Instruct | 1.24B | Q4_K_M | 1.4 GB | 86 | Runs with CPU offload | No |
Fine-tuning
Nothing in the catalogue can be fine-tuned on this machine. CPU fine-tuning is not practical at any useful model size.
Out of reach
25 models will not run here
Qwen2.5 7B Instruct · 7.6BQwen2.5 14B Instruct · 14.8BQwen2.5 32B Instruct · 32.8BQwen2.5 72B Instruct · 72.7BQwen2.5 Coder 7B Instruct · 7.6BQwen2.5 Coder 32B Instruct · 32.8BQwen3 8B · 8.2BQwen3 14B · 14.8BQwen3 30B-A3B · 30.5BLlama 3.1 8B Instruct · 8BLlama 3.3 70B Instruct · 70.6BMistral 7B Instruct v0.3 · 7.25BMistral Nemo 12B Instruct · 12.2BMistral Small 24B Instruct · 23.6BMixtral 8x7B Instruct · 46.7BGemma 2 9B Instruct · 9.2BGemma 3 12B Instruct · 12.2BGemma 3 27B Instruct · 27.4BGemma 3 4B Instruct · 4.3BPhi-4 14B · 14.7BDeepSeek-R1-Distill-Qwen-14B · 14.8BDeepSeek-R1-Distill-Qwen-32B · 32.8BDeepSeek-Coder-V2-Lite Instruct · 15.7BStarCoder2 15B · 16BCode Llama 7B Instruct · 6.7B