Gemma 3 4B Instruct
Vision plus 128k context on a laptop. The most capable genuinely small multimodal option.
- Params
- 4.3B
- Quant
- Q4_K_M
- Memory
- 4.3 GB
Loading…
A 6 GB card has about 5.4 GB usable once the driver and desktop take their share. Every model below is sized against that budget, including the KV cache for a working context.
Vision plus 128k context on a laptop. The most capable genuinely small multimodal option.
Small enough for a laptop CPU or a 6 GB card, and still coherent. A sensible target for edge deployment.
The model that made local AI practical. Still a dependable, fast, unconditionally Apache-licensed baseline.
| # | Model | Params | Licence | Quantization | Memory | Fit | Fine-tune |
|---|---|---|---|---|---|---|---|
| 1 | Gemma 3 4B Instruct | 4.3B | Gemma Terms of Use | Q4_K_M | 4.3 GB | Good fit | Yes |
| 2 | Llama 3.2 3B Instruct | 3.2B | Llama 3.2 Community License | Q4_K_M | 3.2 GB | Excellent fit | Yes |
| 3 | Mistral 7B Instruct v0.3 | 7.25B | Apache 2.0 | Q3_K_M | 4.9 GB | Tight fit | No |
| 4 | Phi-4 14B | 14.7B | MIT | Q4_K_M | 11 GB | Runs with CPU offload | No |
| 5 | Mistral Small 24B Instruct | 23.6B | Apache 2.0 | Q4_K_M | 16 GB | Runs with CPU offload | No |
| 6 | Qwen3 30B-A3B | 30.5B | Apache 2.0 | Q3_K_M | 15 GB | Runs with CPU offload | No |
| 7 | SmolLM2 1.7B Instruct | 1.7B | Apache 2.0 | Q8_0 | 3.7 GB | Good fit | Yes |
| 8 | Qwen2.5 14B Instruct | 14.8B | Apache 2.0 | Q4_K_M | 11 GB | Runs with CPU offload | No |
| 9 | Llama 3.2 1B Instruct | 1.24B | Llama 3.2 Community License | F16 | 3.0 GB | Excellent fit | Yes |
| 10 | StarCoder2 15B | 16B | BigCode OpenRAIL-M | Q4_K_M | 10 GB | Runs with CPU offload | No |
| 11 | Qwen3 14B | 14.8B | Apache 2.0 | Q4_K_M | 10 GB | Runs with CPU offload | No |
| 12 | DeepSeek-R1-Distill-Qwen-14B | 14.8B | MIT | Q4_K_M | 11 GB | Runs with CPU offload | No |
| 13 | Mistral Nemo 12B Instruct | 12.2B | Apache 2.0 | Q4_K_M | 9.1 GB | Runs with CPU offload | No |
| 14 | Qwen2.5 Coder 7B Instruct | 7.6B | Apache 2.0 | Q4_K_M | 5.3 GB | Runs with CPU offload | No |
| 15 | Gemma 3 12B Instruct | 12.2B | Gemma Terms of Use | Q4_K_M | 10 GB | Runs with CPU offload | No |
| 16 | DeepSeek-Coder-V2-Lite Instruct | 15.7B | DeepSeek License | Q4_K_M | 11 GB | Runs with CPU offload | No |
| 17 | Gemma 2 9B Instruct | 9.2B | Gemma Terms of Use | Q4_K_M | 8.1 GB | Runs with CPU offload | No |
| 18 | Qwen2.5 7B Instruct | 7.6B | Apache 2.0 | Q4_K_M | 5.3 GB | Runs with CPU offload | No |
| 19 | Llama 3.1 8B Instruct | 8B | Llama 3.1 Community License | Q3_K_M | 5.2 GB | Runs with CPU offload | No |
| 20 | Qwen3 8B | 8.2B | Apache 2.0 | Q3_K_M | 5.5 GB | Runs with CPU offload | No |
| 21 | Code Llama 7B Instruct | 6.7B | Llama 2 Community License | Q4_K_M | 8.3 GB | Runs with CPU offload | No |
| 22 | Phi-3.5 Mini Instruct | 3.8B | MIT | Q4_K_M | 5.7 GB | Runs with CPU offload | Yes |
Each model is sized at every quantization it publishes: weights at the true bits-per-weight, KV cache for the context window, plus runtime overhead. The recommendation is the highest quality that still leaves room for a real context.
ModelLM can read your actual GPU, VRAM and RAM and size every model against it.
Detect my hardware