Mistral 7B Instruct v0.3
Mistral AI
The model that made local AI practical. Still a dependable, fast, unconditionally Apache-licensed baseline.
Excellent fit
- Weights
- 7.2 GB
- KV cache
- 1.0 GB
- Overhead
- 580 MB
- 8.8 GB of 15 GB usable VRAM.
- Chosen as the best quality that still fits a 32,768-token context (12 GB at that length).
- Apache 2.0 with no usage policy attached
- Very fast
- Extremely well understood
- Outclassed on benchmarks by newer 7B models
- Small vocabulary hurts non-European languages
mistral:7bRunning Mistral 7B Instruct v0.3 on Radeon RX 7800 XT
| Quantization | Quality | Weights | KV cache | Total | ~tok/s | Fit |
|---|---|---|---|---|---|---|
| F16 | lossless | 14 GB | 1.0 GB | 15 GB | 33 | Runs with CPU offload |
| Q8_0Pick | near-lossless | 7.2 GB | 1.0 GB | 8.8 GB | 63 | Excellent fit |
| Q6_K | near-lossless | 5.5 GB | 1.0 GB | 7.1 GB | 81 | Excellent fit |
| Q5_K_M | high | 4.8 GB | 1.0 GB | 6.4 GB | 94 | Excellent fit |
| Q4_K_M | balanced | 4.1 GB | 1.0 GB | 5.7 GB | 110 | Excellent fit |
| Q3_K_M | degraded | 3.3 GB | 1.0 GB | 4.9 GB | 136 | Excellent fit |
KV cache is sized at 8,192 tokens. Longer contexts cost proportionally more — the recommendation above reserves room for a working context.
You can fine-tune this here
- Base weights
- 3.8 GB
- Optimizer
- 625 MB
- Activations
- 900 MB
- Peak
- 6.8 GB
- 6.8 GB peak against 15 GB usable — room to raise batch size or sequence length.
- AMD requires a ROCm build of PyTorch; the Unsloth fast path is CUDA-only.
- Base weights
- 14 GB
- Optimizer
- 625 MB
- Activations
- 900 MB
- Peak
- 17 GB
- Needs 17 GB — switch to QLoRA to cut the weight footprint.
- AMD requires a ROCm build of PyTorch; the Unsloth fast path is CUDA-only.
Where this model runs
VRAM 32 GB · F16 · 15 GB
VRAM 24 GB · F16 · 15 GB
VRAM 16 GB · Q8_0 · 8.8 GB
VRAM 16 GB · Q8_0 · 8.8 GB
VRAM 24 GB · F16 · 15 GB
VRAM 12 GB · Q4_K_M · 5.7 GB
VRAM 12 GB · Q4_K_M · 5.7 GB
VRAM 16 GB · Q8_0 · 8.8 GB
VRAM 48 GB · F16 · 15 GB
VRAM 80 GB · F16 · 15 GB
Unified 128 GB · F16 · 15 GB
Unified 48 GB · F16 · 15 GB
Unified 24 GB · Q8_0 · 8.8 GB
Unified 192 GB · F16 · 15 GB
Unified 16 GB · Q5_K_M · 6.4 GB
VRAM 24 GB · F16 · 15 GB
VRAM 16 GB · Q8_0 · 8.8 GB
VRAM 0 MB
VRAM 0 MB
62.5%
Mistral 7B model card
Reported by the model's author. ModelLM has not run these benchmarks and does not treat them as verified.
Qwen2.5 7B Instruct
7.6B · Apache 2.0
The default starting point for local work on 8–12 GB cards. Strong instruction following and reliable tool-call formatting for its size.
Qwen2.5 Coder 7B Instruct
7.6B · Apache 2.0
The practical local copilot. Supports fill-in-the-middle, so it works as an inline completion model rather than only a chat assistant.
Qwen3 8B
8.2B · Apache 2.0
Switchable thinking mode: the same weights answer directly or reason step by step depending on the prompt. Long context for its size.
Llama 3.1 8B Instruct
8B · Llama 3.1 Community License
The most widely supported open model there is. If a tool, adapter or tutorial exists, it was written for this one first.
Catalogue figures come from each model’s published card. ModelLM has not independently measured them.