Coderhouse goes global

🌍

Read more

What Is LLM Quantization and Why Teams Use INT8 and INT4

AI Engineering

Build and ship real applications with LLMs, agents and automation.

8 weeks · Live online · View course →

Data Analytics

Go from zero to job-ready analyst with SQL, Python and BI tools.

10 weeks · Live online · View course →

Dan Patiño

AI Strategy & Innovation at Coderhouse

Artificial Intelligence

What Is LLM Quantization and Why Teams Use INT8 and INT4

LLM quantization is the practice of representing a large language model’s weights — and sometimes activations — with fewer bits than full floating-point precision, so the same network uses less memory and often runs faster. Instead of storing every parameter as float32 or bfloat16, teams map values into formats such as INT8, INT4, or FP8, accepting a controlled accuracy trade-off in exchange for cheaper inference.

Why it matters: serving frontier-sized models at full precision is expensive, and many agent workloads need lower latency on modest GPUs. Hugging Face’s quantization concept guide and NVIDIA TensorRT-LLM’s quantization docs frame the same goal: shrink the memory footprint and computational cost without retraining a new architecture from scratch.

How Quantization Works (INT8, INT4, and Friends)

Quantization maps a continuous float range into a smaller set of representable values. For INT8, that is typically 256 levels; for INT4, only 16. A scale (and sometimes a zero-point) converts between float and integer space so matrix multiplies can run in low precision and results can be dequantized when needed. INT8 often halves memory versus bf16 with small quality loss; INT4 can roughly quarter memory, but needs stronger methods to stay useful.

Weight-only recipes (for example W4A16 or W8A16) keep activations in higher precision while packing weights tightly — helpful when memory bandwidth is the bottleneck. Activation quantization (W8A8 SmoothQuant, FP8) goes further when hardware accelerates those ops. Methods such as GPTQ and AWQ calibrate on sample data so 4-bit weights retain more task quality than naive rounding. Libraries like bitsandbytes make on-the-fly 8-bit or 4-bit loading practical for experimentation and QLoRA-style fine-tuning.

Quantization vs Distillation vs Smaller Models

Quantization compresses an existing checkpoint numerically. Model distillation trains a smaller student to imitate a teacher’s behavior — a different lever with different failure modes. Teams also ship purpose-built small language models when the task never needed a giant teacher. In production LLMOps, these options often stack: distill or choose an SLM for the control plane, then quantize the serving weights so agents fit on available hardware.

What Production Teams Must Measure

Shipping a quantized checkpoint is not the finish line. Measure task evals (not only perplexity), latency and tokens per second on real hardware, KV-cache memory under long contexts, and failure modes such as math or code regressions that appear only at 4-bit. Prefer calibrated formats (AWQ/GPTQ) for customer-facing agents; use aggressive on-the-fly 4-bit mainly for prototypes. Track which quant recipe each deployment uses so rollbacks and A/B tests stay honest.

PTQ, QAT, and Calibration Choices

Most teams start with post-training quantization (PTQ): take a finished checkpoint, calibrate on a small representative dataset, and emit a lower-bit file. Quantization-aware training (QAT) simulates rounding during training so the model adapts to low precision; it usually recovers more accuracy at aggressive bit-widths but costs more engineering time. In practice, product teams ship PTQ with GPTQ/AWQ for serving, reserve QAT for hardware-constrained devices, and keep a bf16 (or fp8) baseline for debugging.

Calibration data quality matters as much as the algorithm. If your calibration set is only generic web text while production agents do tool calls and Spanish support tickets, 4-bit weights may look fine on public benchmarks and fail on your eval suite. Mirror production prompts, languages, and tool schemas in calibration and in the gates that approve a quantized promotion.

Common Mistakes Teams Make

Teams confuse “it loads on my GPU” with “quality is unchanged.” Others mix quantization and distillation vocabulary in runbooks, so on-call engineers cannot tell whether a regression came from bit-width or from a new student model. Some quantize once offline and never re-check after prompt, tool, or context changes. And many skip hardware fit: FP8 shines on newer accelerators but buys little on older cards. Pair quantization choices with eval gates and cost dashboards before agents depend on them.

If you are building AI products that must stay fast and affordable, structured practice helps. Explore Coderhouse’s AI Engineering course to strengthen model-serving judgment, evaluation habits, and agent architectures that survive real latency budgets.

FAQ

Is INT4 always better than INT8?

No. INT4 saves more memory but usually costs more accuracy. Start with INT8 or a strong calibrated 4-bit recipe, and only go lower when evals and hardware profiling justify it.

Does quantization replace fine-tuning?

No. Quantization changes numeric precision of weights you already have. Fine-tuning (including QLoRA on quantized bases) changes what the model knows. Many stacks do both: adapt the model, then serve a quantized checkpoint.

What is the safest first step?

Load a vendor- or community-calibrated INT8 or AWQ/GPTQ INT4 checkpoint, run your existing agent eval suite, and compare latency plus failure cases against the bf16 baseline before switching production traffic.

Dan Patiño

I'm Dan Patiño, head of AI Strategy & Innovation at Coderhouse. My day-to-day work involves merging the tactical management of e-commerce (CRO, Email Marketing and SEO) with the development of disruptive solutions. I specialize in building internal AI-powered apps to automate tasks and boost innovation within the team. I firmly believe that technology is strategy's best ally. To dive deeper into my professional journey, I'll be waiting for you on my LinkedIn profile.

© 2026 Coderhouse. All rights reserved.

© 2026 Coderhouse. All rights reserved.

© 2026 Coderhouse. All rights reserved.