LearnThatStack Ace your next interview
Open-Source & Local LLMs · question
Question 5 of 55

What is quantization, and why does it matter so much for local LLMs?

beginner
← All Open-Source & Local LLMs questions
Re-explain

Quantization stores model weights, activations, or the key-value attention cache with fewer bits. The attention cache holds information from earlier tokens during generation. Values trained at 16-bit precision may be stored with 8, 4, or fewer bits using scaling factors.

Weights usually dominate the model's memory footprint. Fewer bits reduce required video memory (VRAM) and can speed generation when memory bandwidth is the limit.

Why it matters locally: a 70B model at 16-bit needs around 140 GB, which is multiple data-center GPUs. The same model at 4-bit needs roughly 40 GB and fits on a single high-end card, or even a well-equipped workstation.

The tradeoff is quality. Aggressive quantization introduces rounding error that can degrade accuracy, and small models feel it more than large ones.

Rewriting in plainer words…

This answer doesn't lend itself to a diagram - it reads best . No credits were charged.

Why there's no diagram: “”

The interactive diagram is below the answer - jump to diagram ↓ · Below it, the related concept . Jump to it ↓

The diagram below the answer is the concept . Jump to it ↓

Tailored explanation · switch back to · ·
What should the new diagram focus on?
How well did you know this?
AI:

Saved in this browser - sign in to keep your review list.

How should your speech become text?

Listening… your words appear above as you speak - tap Stop when you're done.

Recording · cr - tap Stop & transcribe when you're done.

Transcribing with AI…

Voice:

Keep going - a few more words and AI can grade it.

Interview lens

Likely follow-ups, what you can say, and the weak answers to avoid.

Sign in free to open it Free account - the lens opens as soon as you're back.

Want a quick review of the fundamentals? See the Open-Source & Local LLMs cheatsheet.

← Back to all Open-Source & Local LLMs questions
Pro · $10/mo

48 of 55 Open-Source & Local LLMs answers are in Pro.

Full answers, code samples, and AI explanations that go simpler or deeper. Cancel anytime.

  • Full answers + code
  • AI explanations, simpler or deeper
  • 1,000 AI credits / month
  • Cancel anytime