Mixtral 8x7B Instruct
Mistral AI
The original open mixture-of-experts. 47B resident, ~13B active — fast generation if you have the memory to hold it.
Mixtral 8x7B Instruct does not fit on CPU only — 64 GB workstation.
CPU-only inference. Expect a few tokens per second at best.
- Generates at ~13B speed
- Apache 2.0
- Well-supported MoE
- Needs ~28 GB even at Q4_K_M
- Superseded on quality by newer dense models
mixtral:8x7bRunning Mixtral 8x7B Instruct on CPU only — 64 GB workstation
| Quantization | Quality | Weights | KV cache | Total | ~tok/s | Fit |
|---|---|---|---|---|---|---|
| F16 | lossless | 87 GB | 1.0 GB | 89 GB | 4 | Won't run |
| Q8_0 | near-lossless | 46 GB | 1.0 GB | 49 GB | 7 | Won't run |
| Q6_K | near-lossless | 36 GB | 1.0 GB | 38 GB | 9 | Not recommended |
| Q5_K_M | high | 31 GB | 1.0 GB | 33 GB | 10 | Not recommended |
| Q4_K_M | balanced | 26 GB | 1.0 GB | 29 GB | 12 | Not recommended |
| Q3_K_M | degraded | 21 GB | 1.0 GB | 24 GB | 15 | Not recommended |
KV cache is sized at 8,192 tokens. Longer contexts cost proportionally more — the recommendation above reserves room for a working context.
Local fine-tuning on this machine
- Base weights
- 24 GB
- Optimizer
- 313 MB
- Activations
- 900 MB
- Peak
- 28 GB
- Local fine-tuning needs a GPU or Apple Silicon. CPU training is not practical.
- Base weights
- 87 GB
- Optimizer
- 313 MB
- Activations
- 900 MB
- Peak
- 90 GB
- Local fine-tuning needs a GPU or Apple Silicon. CPU training is not practical.
Where this model runs
VRAM 32 GB · Q4_K_M · 29 GB
VRAM 24 GB · Q4_K_M · 29 GB
VRAM 16 GB · Q4_K_M · 29 GB
VRAM 16 GB · Q4_K_M · 29 GB
VRAM 24 GB · Q4_K_M · 29 GB
VRAM 12 GB · Q4_K_M · 29 GB
VRAM 12 GB
VRAM 16 GB · Q4_K_M · 29 GB
VRAM 48 GB · Q5_K_M · 33 GB
VRAM 80 GB · Q8_0 · 49 GB
Unified 128 GB · Q8_0 · 49 GB
Unified 48 GB · Q4_K_M · 29 GB
Unified 24 GB · Q4_K_M · 29 GB
Unified 192 GB · F16 · 89 GB
Unified 16 GB · Q3_K_M · 24 GB
VRAM 24 GB · Q4_K_M · 29 GB
VRAM 16 GB · Q4_K_M · 29 GB
VRAM 0 MB
VRAM 0 MB
70.6%
Mixtral paper
Reported by the model's author. ModelLM has not run these benchmarks and does not treat them as verified.