Nomic Embed Text v1.5
Nomic AI
The default local embedding model for Knowledge mode. 8k context and Matryoshka truncation, so you can trade index size for accuracy.
Runs with CPU offload
- Weights
- 140 MB
- KV cache
- 280 MB
- Overhead
- 450 MB
- CPU-only inference. Expect a few tokens per second at best.
- Small models remain usable on CPU for non-interactive work.
- KV cache at 8,192 tokens is 280 MB — shorten the context to reclaim memory.
- 8192-token context
- Variable output dimensions
- Runs fast on CPU
- English-centric
- Not a generative model
nomic-embed-textRunning Nomic Embed Text v1.5 on CPU only — 64 GB workstation
| Quantization | Quality | Weights | KV cache | Total | ~tok/s | Fit |
|---|---|---|---|---|---|---|
| F16 | lossless | 260 MB | 280 MB | 990 MB | 339 | Runs with CPU offload |
| Q8_0Pick | near-lossless | 140 MB | 280 MB | 870 MB | 637 | Runs with CPU offload |
KV cache is sized at 8,192 tokens. Longer contexts cost proportionally more — the recommendation above reserves room for a working context.
Local fine-tuning on this machine
This model does not publish weights that can be fine-tuned locally.
Where this model runs
VRAM 32 GB · F16 · 990 MB
VRAM 24 GB · F16 · 990 MB
VRAM 16 GB · F16 · 990 MB
VRAM 16 GB · F16 · 990 MB
VRAM 24 GB · F16 · 990 MB
VRAM 12 GB · F16 · 990 MB
VRAM 12 GB · F16 · 990 MB
VRAM 16 GB · F16 · 990 MB
VRAM 48 GB · F16 · 990 MB
VRAM 80 GB · F16 · 990 MB
Unified 128 GB · F16 · 990 MB
Unified 48 GB · F16 · 990 MB
Unified 24 GB · F16 · 990 MB
Unified 192 GB · F16 · 990 MB
Unified 16 GB · F16 · 990 MB
VRAM 24 GB · F16 · 990 MB
VRAM 16 GB · F16 · 990 MB
VRAM 0 MB · Q8_0 · 870 MB
VRAM 0 MB · Q8_0 · 870 MB
Llama 3.2 3B Instruct
3.2B · Llama 3.2 Community License
Small enough for a laptop CPU or a 6 GB card, and still coherent. A sensible target for edge deployment.
Llama 3.2 1B Instruct
1.24B · Llama 3.2 Community License
A genuinely tiny model for classification, routing and extraction rather than open-ended chat.
Gemma 3 4B Instruct
4.3B · Gemma Terms of Use
Vision plus 128k context on a laptop. The most capable genuinely small multimodal option.
Phi-3.5 Mini Instruct
3.8B · MIT
Strong reasoning at 3.8B with a 128k context — but full multi-head attention makes its KV cache expensive.
Catalogue figures come from each model’s published card. ModelLM has not independently measured them.