Best local AI models for CPU only — 64 GB workstation
Enough memory to load large models; not enough bandwidth to generate with them at a usable rate.
Specification
VRAM64 GB
Usable for models45 GB
System RAM64 GB
Bandwidth120 GB/s
Local trainingNot practical
Operating systemswindows, linux
Runs well
4 models fit this machine
Estimated
| Model | Params | Quantization | Memory | ~tok/s | Fit | Fine-tune |
|---|---|---|---|---|---|---|
| Phi-3.5 Mini Instruct | 3.8B | Q4_K_M | 5.7 GB | 40 | Runs with CPU offload | No |
| Llama 3.2 3B Instruct | 3.2B | Q4_K_M | 3.2 GB | 48 | Runs with CPU offload | No |
| SmolLM2 1.7B Instruct | 1.7B | Q4_K_M | 2.9 GB | 90 | Runs with CPU offload | No |
| Llama 3.2 1B Instruct | 1.24B | Q4_K_M | 1.4 GB | 124 | Runs with CPU offload | No |
Fine-tuning
Nothing in the catalogue can be fine-tuned on this machine. Local training needs an accelerator with enough memory to hold the base weights, the adapter, the optimizer state and the activations at once.
Out of reach
25 models will not run here
Qwen2.5 7B Instruct · 7.6BQwen2.5 14B Instruct · 14.8BQwen2.5 32B Instruct · 32.8BQwen2.5 72B Instruct · 72.7BQwen2.5 Coder 7B Instruct · 7.6BQwen2.5 Coder 32B Instruct · 32.8BQwen3 8B · 8.2BQwen3 14B · 14.8BQwen3 30B-A3B · 30.5BLlama 3.1 8B Instruct · 8BLlama 3.3 70B Instruct · 70.6BMistral 7B Instruct v0.3 · 7.25BMistral Nemo 12B Instruct · 12.2BMistral Small 24B Instruct · 23.6BMixtral 8x7B Instruct · 46.7BGemma 2 9B Instruct · 9.2BGemma 3 12B Instruct · 12.2BGemma 3 27B Instruct · 27.4BGemma 3 4B Instruct · 4.3BPhi-4 14B · 14.7BDeepSeek-R1-Distill-Qwen-14B · 14.8BDeepSeek-R1-Distill-Qwen-32B · 32.8BDeepSeek-Coder-V2-Lite Instruct · 15.7BStarCoder2 15B · 16BCode Llama 7B Instruct · 6.7B