GGUF is a model file format used by llama.cpp, Ollama, and the LM Studio app. One file can hold quantized weights, the tokenizer, the chat template, and other model details.
It replaced the older GGML format and is widely used for laptop and CPU inference.
GGUF comes in many quantization levels, and the naming tells you the tradeoff:
Q2andQ3: smallest files, with clear quality loss.Q4: a common balance of size and quality.Q5andQ6: larger files with better quality.Q8: close to full precision but much larger.
# Illustrative: tags vary by model library
ollama pull <model>:<quantized-tag>
GGUF can run on a CPU, GPU, Apple Metal, or a mix. You can place some layers on the GPU and keep the rest in system memory. For high-throughput GPU servers, formats designed for GPU kernels may perform better.
This answer doesn't lend itself to a diagram - it reads best . No credits were charged.
Why there's no diagram: “”
The interactive diagram is below the answer - jump to diagram ↓ · Below it, the related concept . Jump to it ↓
The diagram below the answer is the concept . Jump to it ↓