Llama 3.2 1B Instruct
Meta
A genuinely tiny model for classification, routing and extraction rather than open-ended chat.
Excellent fit
- Weights
- 2.3 GB
- KV cache
- 250 MB
- Overhead
- 470 MB
- 3.0 GB of 47 GB usable VRAM.
- Chosen as the best quality that still fits a 32,768-token context (3.8 GB at that length).
- Runs on a phone or Raspberry Pi
- Fine-tunes in minutes
- Excellent for narrow tasks
- Not usable as a general assistant
- Needs fine-tuning to be useful
llama3.2:1bRunning Llama 3.2 1B Instruct on RTX 6000 Ada Generation
| Quantization | Quality | Weights | KV cache | Total | ~tok/s | Fit |
|---|---|---|---|---|---|---|
| F16Pick | lossless | 2.3 GB | 250 MB | 3.0 GB | 299 | Excellent fit |
| Q8_0 | near-lossless | 1.2 GB | 250 MB | 1.9 GB | 563 | Excellent fit |
| Q6_K | near-lossless | 950 MB | 250 MB | 1.7 GB | 730 | Excellent fit |
| Q5_K_M | high | 820 MB | 250 MB | 1.5 GB | 844 | Excellent fit |
| Q4_K_M | balanced | 700 MB | 250 MB | 1.4 GB | 991 | Excellent fit |
| Q3_K_M | degraded | 560 MB | 250 MB | 1.3 GB | 1225 | Excellent fit |
KV cache is sized at 8,192 tokens. Longer contexts cost proportionally more — the recommendation above reserves room for a working context.
You can fine-tune this here
- Base weights
- 650 MB
- Optimizer
- 168 MB
- Activations
- 340 MB
- Peak
- 2.3 GB
- 2.3 GB peak against 47 GB usable — room to raise batch size or sequence length.
- Base weights
- 2.3 GB
- Optimizer
- 168 MB
- Activations
- 340 MB
- Peak
- 4.0 GB
- 4.0 GB peak against 47 GB usable — room to raise batch size or sequence length.
Where this model runs
VRAM 32 GB · F16 · 3.0 GB
VRAM 24 GB · F16 · 3.0 GB
VRAM 16 GB · F16 · 3.0 GB
VRAM 16 GB · F16 · 3.0 GB
VRAM 24 GB · F16 · 3.0 GB
VRAM 12 GB · F16 · 3.0 GB
VRAM 12 GB · F16 · 3.0 GB
VRAM 16 GB · F16 · 3.0 GB
VRAM 48 GB · F16 · 3.0 GB
VRAM 80 GB · F16 · 3.0 GB
Unified 128 GB · F16 · 3.0 GB
Unified 48 GB · F16 · 3.0 GB
Unified 24 GB · F16 · 3.0 GB
Unified 192 GB · F16 · 3.0 GB
Unified 16 GB · F16 · 3.0 GB
VRAM 24 GB · F16 · 3.0 GB
VRAM 16 GB · F16 · 3.0 GB
VRAM 0 MB · Q4_K_M · 1.4 GB
VRAM 0 MB · Q4_K_M · 1.4 GB
Llama 3.1 8B Instruct
8B · Llama 3.1 Community License
The most widely supported open model there is. If a tool, adapter or tutorial exists, it was written for this one first.
Llama 3.2 3B Instruct
3.2B · Llama 3.2 Community License
Small enough for a laptop CPU or a 6 GB card, and still coherent. A sensible target for edge deployment.
Llama 3.3 70B Instruct
70.6B · Llama 3.3 Community License
Delivers close to Llama 3.1 405B quality at a size a dual-GPU workstation or 64 GB Mac can actually hold.
Gemma 3 4B Instruct
4.3B · Gemma Terms of Use
Vision plus 128k context on a laptop. The most capable genuinely small multimodal option.
Catalogue figures come from each model’s published card. ModelLM has not independently measured them.