Best overallExcellent fit
Qwen2.5 Coder 32B Instruct
The strongest open code model that still fits one 24 GB card. The reason a lot of people buy a 4090.
- Params
- 32.8B
- Quant
- Q4_K_M
- Memory
- 21 GB
Loading…
Holds a 32B at Q4_K_M with room for context. Bandwidth, not capacity, is the limit here.
The strongest open code model that still fits one 24 GB card. The reason a lot of people buy a 4090.
Approaches 70B quality at half the memory. Runs on a single 24 GB card at Q4_K_M with a modest context window.
The strongest open reasoning model that fits a single 24 GB card. A genuinely different capability class on hard problems.
| # | Model | Params | Licence | Quantization | Memory | ~tok/s | Fit | Fine-tune |
|---|---|---|---|---|---|---|---|---|
| 1 | Qwen2.5 Coder 32B Instruct | 32.8B | Apache 2.0 | Q4_K_M | 21 GB | 11 | Excellent fit | Yes |
| 2 | Qwen2.5 32B Instruct | 32.8B | Apache 2.0 | Q4_K_M | 21 GB | 11 | Excellent fit | Yes |
| 3 | DeepSeek-R1-Distill-Qwen-32B | 32.8B | MIT | Q4_K_M | 21 GB | 11 | Excellent fit | Yes |
| 4 | Mixtral 8x7B Instruct | 46.7B | Apache 2.0 | Q4_K_M | 29 GB | 27 | Good fit | Yes |
| 5 | Gemma 3 27B Instruct | 27.4B | Gemma Terms of Use | Q4_K_M | 21 GB | 13 | Excellent fit | Yes |
| 6 | Phi-4 14B | 14.7B | MIT | Q8_0 | 17 GB | 14 | Excellent fit | Yes |
| 7 | Mistral Small 24B Instruct | 23.6B | Apache 2.0 | Q6_K | 20 GB | 11 | Excellent fit | Yes |
| 8 | Qwen2.5 14B Instruct | 14.8B | Apache 2.0 | Q8_0 | 17 GB | 13 | Excellent fit | Yes |
| 9 | Qwen3 30B-A3B | 30.5B | Apache 2.0 | Q6_K | 25 GB | 78 | Good fit | Yes |
| 10 | StarCoder2 15B | 16B | BigCode OpenRAIL-M | Q8_0 | 17 GB | 12 | Excellent fit | Yes |
| 11 | Qwen3 14B | 14.8B | Apache 2.0 | Q8_0 | 17 GB | 13 | Excellent fit | Yes |
| 12 | DeepSeek-R1-Distill-Qwen-14B | 14.8B | MIT | Q8_0 | 17 GB | 13 | Excellent fit | Yes |
| 13 | Mistral Nemo 12B Instruct | 12.2B | Apache 2.0 | Q8_0 | 14 GB | 16 | Excellent fit | Yes |
| 14 | Qwen2.5 Coder 7B Instruct | 7.6B | Apache 2.0 | F16 | 15 GB | 14 | Excellent fit | Yes |
| 15 | Gemma 3 12B Instruct | 12.2B | Gemma Terms of Use | Q8_0 | 16 GB | 16 | Excellent fit | Yes |
| 16 | DeepSeek-Coder-V2-Lite Instruct | 15.7B | DeepSeek License | Q8_0 | 18 GB | 83 | Excellent fit | Yes |
| 17 | Gemma 2 9B Instruct | 9.2B | Gemma Terms of Use | F16 | 20 GB | 11 | Excellent fit | Yes |
| 18 | Qwen2.5 7B Instruct | 7.6B | Apache 2.0 | F16 | 15 GB | 14 | Excellent fit | Yes |
| 19 | Llama 3.1 8B Instruct | 8B | Llama 3.1 Community License | F16 | 16 GB | 13 | Excellent fit | Yes |
| 20 | Qwen3 8B | 8.2B | Apache 2.0 | F16 | 17 GB | 13 | Excellent fit | Yes |
| 21 | Mistral 7B Instruct v0.3 | 7.25B | Apache 2.0 | F16 | 15 GB | 15 | Excellent fit | Yes |
| 22 | Code Llama 7B Instruct | 6.7B | Llama 2 Community License | F16 | 17 GB | 16 | Excellent fit | Yes |
| 23 | Phi-3.5 Mini Instruct | 3.8B | MIT | F16 | 11 GB | 28 | Excellent fit | Yes |
| 24 | Gemma 3 4B Instruct | 4.3B | Gemma Terms of Use | F16 | 9.9 GB | 25 | Excellent fit | Yes |
| 25 | Llama 3.2 3B Instruct | 3.2B | Llama 3.2 Community License | F16 | 7.3 GB | 33 | Excellent fit | Yes |
| 26 | Qwen2.5 72B Instruct | 72.7B | Qwen License | Q4_K_M | 45 GB | 5 | Runs with CPU offload | No |
| 27 | Llama 3.3 70B Instruct | 70.6B | Llama 3.3 Community License | Q4_K_M | 44 GB | 5 | Runs with CPU offload | No |
| 28 | SmolLM2 1.7B Instruct | 1.7B | Apache 2.0 | F16 | 5.2 GB | 62 | Excellent fit | Yes |
| 29 | Llama 3.2 1B Instruct | 1.24B | Llama 3.2 Community License | F16 | 3.0 GB | 85 | Excellent fit | Yes |
Memory is calculated from each model’s published geometry against this machine’s usable capacity. Where the vendor publishes a memory bandwidth figure, an order-of-magnitude token rate is derived from it and labelled as an estimate.
ModelLM can read your actual GPU, VRAM and RAM and size every model against it.
Detect my hardware