Mistral Small 24B Instruct
Mistral AI
Built deliberately for low latency on a single card — fewer layers, wider FFN. Apache 2.0 at a size that usually is not.
Good fit
- Weights
- 13 GB
- KV cache
- 1.6 GB
- Overhead
- 870 MB
- 16 GB of 23 GB usable VRAM.
- Chosen as the best quality that still fits a 16,384-token context (17 GB at that length).
- Low-latency architecture
- Apache 2.0 at 24B
- Strong function calling
- 32k context, short by 2025 standards
- Q4 required on 24 GB cards
mistral-small:24bRunning Mistral Small 24B Instruct on GeForce RTX 4090
| Quantization | Quality | Weights | KV cache | Total | ~tok/s | Fit |
|---|---|---|---|---|---|---|
| F16 | lossless | 44 GB | 1.6 GB | 46 GB | 17 | Runs with CPU offload |
| Q8_0 | near-lossless | 23 GB | 1.6 GB | 26 GB | 31 | Runs with CPU offload |
| Q6_K | near-lossless | 18 GB | 1.6 GB | 20 GB | 40 | Tight fit |
| Q5_K_M | high | 16 GB | 1.6 GB | 18 GB | 47 | Good fit |
| Q4_K_MPick | balanced | 13 GB | 1.6 GB | 16 GB | 55 | Good fit |
| Q3_K_M | degraded | 11 GB | 1.6 GB | 13 GB | 68 | Excellent fit |
KV cache is sized at 8,192 tokens. Longer contexts cost proportionally more — the recommendation above reserves room for a working context.
You can fine-tune this here
- Base weights
- 12 GB
- Optimizer
- 701 MB
- Activations
- 1.7 GB
- Peak
- 17 GB
- 17 GB peak against 23 GB usable.
- Base weights
- 44 GB
- Optimizer
- 701 MB
- Activations
- 1.7 GB
- Peak
- 48 GB
- Needs 48 GB — switch to QLoRA to cut the weight footprint.
Where this model runs
VRAM 32 GB · Q5_K_M · 18 GB
VRAM 24 GB · Q4_K_M · 16 GB
VRAM 16 GB · Q3_K_M · 13 GB
VRAM 16 GB · Q3_K_M · 13 GB
VRAM 24 GB · Q4_K_M · 16 GB
VRAM 12 GB · Q4_K_M · 16 GB
VRAM 12 GB · Q4_K_M · 16 GB
VRAM 16 GB · Q3_K_M · 13 GB
VRAM 48 GB · Q8_0 · 26 GB
VRAM 80 GB · F16 · 46 GB
Unified 128 GB · F16 · 46 GB
Unified 48 GB · Q6_K · 20 GB
Unified 24 GB · Q4_K_M · 16 GB
Unified 192 GB · F16 · 46 GB
Unified 16 GB · Q4_K_M · 16 GB
VRAM 24 GB · Q4_K_M · 16 GB
VRAM 16 GB · Q3_K_M · 13 GB
VRAM 0 MB
VRAM 0 MB
Mistral 7B Instruct v0.3
7.25B · Apache 2.0
The model that made local AI practical. Still a dependable, fast, unconditionally Apache-licensed baseline.
Mistral Nemo 12B Instruct
12.2B · Apache 2.0
A 12B with a 128k context and the Tekken tokenizer, which compresses non-English text far better than Llama’s.
Gemma 3 27B Instruct
27.4B · Gemma Terms of Use
The largest Gemma 3. Multimodal, long-context, and designed to run on a single high-memory accelerator.
Catalogue figures come from each model’s published card. ModelLM has not independently measured them.