BGE-M3
BAAI
Multilingual retrieval across 100+ languages, with dense, sparse and multi-vector output from one model.
Excellent fit
- Weights
- 1.1 GB
- KV cache
- 750 MB
- Overhead
- 460 MB
- 2.3 GB of 23 GB usable VRAM.
- KV cache at 8,192 tokens is 750 MB — shorten the context to reclaim memory.
- Chosen as the best quality that still fits a 8,192-token context (2.3 GB at that length).
- Genuinely multilingual retrieval
- Hybrid dense + sparse scoring
- 8k context
- Heavier than Nomic Embed
- Large vocabulary increases memory
bge-m3Running BGE-M3 on Radeon RX 7900 XTX
| Quantization | Quality | Weights | KV cache | Total | ~tok/s | Fit |
|---|---|---|---|---|---|---|
| F16Pick | lossless | 1.1 GB | 750 MB | 2.3 GB | 653 | Excellent fit |
| Q8_0 | near-lossless | 560 MB | 750 MB | 1.8 GB | 1230 | Excellent fit |
KV cache is sized at 8,192 tokens. Longer contexts cost proportionally more — the recommendation above reserves room for a working context.
Local fine-tuning on this machine
This model does not publish weights that can be fine-tuned locally.
Where this model runs
VRAM 32 GB · F16 · 2.3 GB
VRAM 24 GB · F16 · 2.3 GB
VRAM 16 GB · F16 · 2.3 GB
VRAM 16 GB · F16 · 2.3 GB
VRAM 24 GB · F16 · 2.3 GB
VRAM 12 GB · F16 · 2.3 GB
VRAM 12 GB · F16 · 2.3 GB
VRAM 16 GB · F16 · 2.3 GB
VRAM 48 GB · F16 · 2.3 GB
VRAM 80 GB · F16 · 2.3 GB
Unified 128 GB · F16 · 2.3 GB
Unified 48 GB · F16 · 2.3 GB
Unified 24 GB · F16 · 2.3 GB
Unified 192 GB · F16 · 2.3 GB
Unified 16 GB · F16 · 2.3 GB
VRAM 24 GB · F16 · 2.3 GB
VRAM 16 GB · F16 · 2.3 GB
VRAM 0 MB · Q8_0 · 1.8 GB
VRAM 0 MB · Q8_0 · 1.8 GB
Llama 3.2 3B Instruct
3.2B · Llama 3.2 Community License
Small enough for a laptop CPU or a 6 GB card, and still coherent. A sensible target for edge deployment.
Llama 3.2 1B Instruct
1.24B · Llama 3.2 Community License
A genuinely tiny model for classification, routing and extraction rather than open-ended chat.
Gemma 3 4B Instruct
4.3B · Gemma Terms of Use
Vision plus 128k context on a laptop. The most capable genuinely small multimodal option.
Phi-3.5 Mini Instruct
3.8B · MIT
Strong reasoning at 3.8B with a 128k context — but full multi-head attention makes its KV cache expensive.
Catalogue figures come from each model’s published card. ModelLM has not independently measured them.