Llama 3.3 70B Instruct
Meta
Delivers close to Llama 3.1 405B quality at a size a dual-GPU workstation or 64 GB Mac can actually hold.
Runs with CPU offload
- Weights
- 32 GB
- KV cache
- 2.5 GB
- Overhead
- 1.6 GB
- Exceeds VRAM by 21 GB; layers spill to system RAM and generation slows sharply.
- Top-tier open general quality
- 128k context
- Broad tooling support
- Needs 48 GB+ for comfortable use
- Local fine-tuning is a multi-GPU exercise
llama3.3:70bRunning Llama 3.3 70B Instruct on GeForce RTX 5080
| Quantization | Quality | Weights | KV cache | Total | ~tok/s | Fit |
|---|---|---|---|---|---|---|
| F16 | lossless | 132 GB | 2.5 GB | 136 GB | 5 | Won't run |
| Q8_0 | near-lossless | 70 GB | 2.5 GB | 74 GB | 10 | Won't run |
| Q6_K | near-lossless | 54 GB | 2.5 GB | 58 GB | 13 | Won't run |
| Q5_K_M | high | 47 GB | 2.5 GB | 51 GB | 15 | Won't run |
| Q4_K_M | balanced | 40 GB | 2.5 GB | 44 GB | 17 | Won't run |
| Q3_K_MPick | degraded | 32 GB | 2.5 GB | 36 GB | 22 | Runs with CPU offload |
KV cache is sized at 8,192 tokens. Longer contexts cost proportionally more — the recommendation above reserves room for a working context.
Local fine-tuning on this machine
- Base weights
- 37 GB
- Optimizer
- 1.5 GB
- Activations
- 2.9 GB
- Peak
- 44 GB
- QLoRA still needs 44 GB; this machine has 15 GB. Choose a smaller base model.
- Base weights
- 132 GB
- Optimizer
- 1.5 GB
- Activations
- 2.9 GB
- Peak
- 139 GB
- Needs 139 GB — switch to QLoRA to cut the weight footprint.
Where this model runs
VRAM 32 GB · Q4_K_M · 44 GB
VRAM 24 GB · Q4_K_M · 44 GB
VRAM 16 GB · Q3_K_M · 36 GB
VRAM 16 GB · Q3_K_M · 36 GB
VRAM 24 GB · Q4_K_M · 44 GB
VRAM 12 GB
VRAM 12 GB
VRAM 16 GB · Q3_K_M · 36 GB
VRAM 48 GB · Q4_K_M · 44 GB
VRAM 80 GB · Q5_K_M · 51 GB
Unified 128 GB · Q6_K · 58 GB
Unified 48 GB · Q4_K_M · 44 GB
Unified 24 GB
Unified 192 GB · Q8_0 · 74 GB
Unified 16 GB
VRAM 24 GB · Q4_K_M · 44 GB
VRAM 16 GB · Q3_K_M · 36 GB
VRAM 0 MB
VRAM 0 MB
86%
Llama 3.3 model card
88.4%
Llama 3.3 model card
Reported by the model's author. ModelLM has not run these benchmarks and does not treat them as verified.
Qwen2.5 72B Instruct
72.7B · Qwen License
Frontier-adjacent open weights. Needs a workstation, a multi-GPU rig or a large unified-memory Mac.
Llama 3.1 8B Instruct
8B · Llama 3.1 Community License
The most widely supported open model there is. If a tool, adapter or tutorial exists, it was written for this one first.
Llama 3.2 3B Instruct
3.2B · Llama 3.2 Community License
Small enough for a laptop CPU or a 6 GB card, and still coherent. A sensible target for edge deployment.
Llama 3.2 1B Instruct
1.24B · Llama 3.2 Community License
A genuinely tiny model for classification, routing and extraction rather than open-ended chat.
Catalogue figures come from each model’s published card. ModelLM has not independently measured them.