Engineering8 min read

INT4 Quantization for LLMs: How to Measure the Quality Loss Before You Ship

A practical guide to INT4 and FP4 LLM quantization: how GPTQ, AWQ, FP8, NVFP4 and KV cache quantization work, why perplexity hides regressions, and how to build an evaluation plan that tells you whether a quantized model is safe to ship.

Quantization is the cheapest large speedup available for LLM inference. Converting weights from 16-bit to 4-bit cuts model memory by roughly four times and can make decode several times faster on the same GPUs. It also changes the model, and the change rarely shows up where teams look first. A quantized model can match its baseline on a perplexity benchmark and still fail on long documents, multi-step reasoning or a language that was thin in the calibration data.

Our DeepSeek R1 inference world record, 303 tokens per second, ran in FP4, a 4-bit floating-point format, and Artificial Analysis reported that it maintained accuracy against the native FP8 version on their evaluation suite. Whatever the 4-bit format, INT4 or FP4, the number that decides whether it can ship is the quality loss measured on tasks defined before touching the weights. This article describes how to run that measurement on your own model.

What weight quantization does and why it speeds up decode

An LLM generates text one token at a time. For every new token, the GPU reads every active weight from memory and multiplies it by a small activation vector. At typical batch sizes, the arithmetic is cheap and the memory read dominates, so decode is memory-bandwidth bound. Tokens per second are roughly set by how fast the GPU can stream the weights through its compute units.

Weight quantization stores each weight in fewer bits. A 4-bit weight is a quarter of the size of a BF16 weight, so the GPU moves a quarter of the bytes per token. Kernels unpack the 4-bit values and apply scale factors on the fly, usually per group of 32–128 weights, before the matrix multiply. When the kernel is well written, the unpacking cost hides behind the memory read and the speedup tracks the reduction in bytes.

Prefill, where the model processes the prompt, is closer to compute bound, so weight-only quantization helps it less. Speeding up prefill requires quantizing activations too, so the matrix multiply itself runs in a lower-precision format on the tensor cores.

INT4 quantization methods: RTN, GPTQ and AWQ

Round-to-nearest (RTN)

The baseline method: pick a scale per group, round each weight to the nearest representable value, done. RTN needs no calibration data and runs in minutes. At 8 bits it is usually fine. At 4 bits it often costs measurable quality, because a few large weights stretch the scale and every small weight in the group loses resolution.

GPTQ

GPTQ quantizes one layer at a time and uses approximate second-order information, computed from a small calibration set, to correct for rounding error. After quantizing a column of weights, it adjusts the remaining unquantized weights to compensate. It is a one-shot method with no retraining, and it generally recovers much of the gap between RTN and full precision at 4 bits. Its weak point is that it fits the correction to whatever calibration data you feed it.

AWQ

Activation-aware Weight Quantization starts from the observation that a small fraction of weight channels matter far more than the rest, and that you can find them by looking at activation magnitudes. AWQ scales those salient channels up before quantization and folds the inverse scale into the preceding operation, which protects them without mixed precision. It also uses calibration data, though it tends to be less sensitive to its exact choice than GPTQ.

Activation and low-precision float formats: SmoothQuant, FP8 and NVFP4

SmoothQuant for activations

Activations are harder to quantize than weights because a few channels carry outliers that can be orders of magnitude larger than the rest. SmoothQuant moves that difficulty from activations into weights with a per-channel scaling transform that is mathematically equivalent in full precision. The result is that both weights and activations fit in INT8 (W8A8), which lets prefill use integer tensor cores.

FP8

FP8 (E4M3 and E5M2) is supported natively on NVIDIA Hopper and later GPUs. With per-tensor or finer-grained scaling, FP8 weights and activations are close to lossless for many models, and FP8 is often the safest first step for any deployment on that hardware. DeepSeek V3 was trained with FP8 mixed precision, and both V3 and R1 were released with FP8 weights, so for those models 4-bit quantization starts from FP8 rather than BF16.

FP4 and NVFP4

NVIDIA's Blackwell generation adds hardware support for 4-bit floating point. NVFP4 stores values in an E2M1 format with a shared FP8 scale for each block of 16 values, plus a per-tensor scale. The small block size handles outliers better than coarse INT4 groups, and the tensor cores consume the format directly, so both weights and activations can run in 4 bits. The caveat is hardware: on GPUs without native FP4 support, these formats fall back to slower paths or are not supported at all.

KV cache quantization

On long contexts the KV cache can outgrow the weights. Storing keys and values in FP8, or in lower precision with careful scaling, roughly halves or quarters that memory and lets you fit longer contexts or larger batches. Errors here behave differently from weight errors: they accumulate over sequence length and show up as degraded retrieval deep in the context. Test KV cache quantization separately from weight quantization, so you know which change caused a regression.

Why perplexity alone misleads

Perplexity on a generic text corpus is the most common quantization metric and the least informative one for a production decision. Three problems come up repeatedly.

First, perplexity averages over every token. Quantization damage concentrates in rare, high-stakes tokens: a digit in an account number, a closing bracket, a negation. A model can lose those and keep almost the same average.

Second, perplexity scores text the model reads. Production quality depends on text the model writes, often over thousands of tokens, where small errors compound.

Third, generic corpora do not match your workload. Perplexity on English web text says little about Arabic legal documents or long financial filings.

Use perplexity as a smoke test. A large jump means something is broken. A small change means nothing yet.

A practical evaluation plan for quantized LLMs

Build the eval set before you quantize

Write down what the model is for and build a task-specific evaluation set first, with a scoring method for each task: exact match, unit tests for code, rubric grading for long-form answers, or retrieval accuracy for long-context lookups. Building the set after you see quantized outputs invites tuning the test to the result. A few hundred well-chosen examples per task beat a public leaderboard.

Compare against a full-precision baseline on the same stack

Run the unquantized model through the same serving engine, prompts and scorer. The comparison you want is the difference between the two models on identical inputs, including per-example disagreements as well as averages.

Hold sampling settings fixed

Temperature, top-p, max tokens, chat template and system prompt must be identical across runs. Greedy decoding gives the cleanest comparison. If production uses sampling, also run several seeds per prompt and compare distributions, because sampling noise can be larger than the quantization effect.

Cover the regressions perplexity hides

Quantization damage is uneven. Include slices that stress it:

  • Long context: needle-in-a-haystack and multi-document questions at the context lengths you actually serve, with and without KV cache quantization.
  • Reasoning: multi-step math and logic problems where one wrong intermediate step fails the answer.
  • Code: generation scored by execution against tests.
  • Multilingual: every language in your traffic. For deployments in the Gulf, that means Arabic, including dialectal and mixed Arabic-English text, which is often underrepresented in calibration sets.
  • Structured output: JSON and tool calls that must parse.

Choose the calibration set deliberately

GPTQ and AWQ fit to their calibration data. Default calibration sets are mostly English web text. If your workload is code, Arabic or long documents, include representative samples of each, and keep them separate from the eval set so you are not testing on what you calibrated on.

Find the outlier layers

Not every layer tolerates 4 bits equally. Quantize layer by layer, or measure per-layer reconstruction error, to find the sensitive ones. Common candidates are the first and last layers, attention output projections, and in mixture-of-experts models, the router and any shared experts. Keeping a small number of sensitive layers at 8 bits often recovers most of the lost quality at a small memory cost.

Deciding what quality loss is acceptable

"Under 1%" is only meaningful relative to a task. Translate the regression into business terms before deciding:

  • How many more wrong answers per thousand requests, and what does each cost to correct?
  • Does the regression fall on a slice that matters, such as a regulated workflow, a key language or the longest documents?
  • What does the speedup buy: fewer GPUs, lower latency at the same hardware, or capacity for a new use case?

A 2% drop on casual chat may be an easy trade for halving the GPU fleet. A 0.5% drop on contract clause extraction may not be. Set the threshold per task with the people who own the outcome, and write it down before you look at the results.

Deployment checks before you ship

Confirm kernel support for the format

A quantization format is only as fast as the kernel that runs it. Check that your serving engine has optimized kernels for your exact combination of format, group size, GPU architecture and model architecture, including mixture-of-experts layers. A format without a fast kernel can run slower than the unquantized model. When no existing kernel fits, writing one is often where the real speedup comes from; a custom TensorRT-LLM build was a core part of our DeepSeek R1 FP4 work on NVIDIA Blackwell. We cover that side in more detail in how we reached 303 tokens per second on DeepSeek R1 in FP4.

Measure throughput on real traffic shapes

Fixed 128-token benchmarks hide most of what matters. Replay a sample of production traffic, or synthesize one with the same distribution of prompt lengths, output lengths and concurrency. Report time to first token, inter-token latency and total throughput at the batch sizes you will run. Quantization gains are largest at small batches, where decode is most memory bound, and shrink as batches grow and the workload becomes compute bound.

Re-run the eval on the deployed stack

Run the full evaluation one more time through the production engine, with production settings, after integration. Fused operations and different attention backends can shift outputs slightly, and the model you evaluated should be the model you ship.

Summary

Build a task-specific eval set first, compare against a full-precision baseline with fixed settings, stress the slices perplexity hides, and decide the acceptable loss in business terms. Then confirm the kernels exist and measure on real traffic. The same discipline applies whether the target is INT4 on current hardware or FP4 on GPUs that support it natively, as in our DeepSeek R1 record.

If you are planning to quantize a model for production and want the evaluation done properly, see our LLM engineering services or get in touch.

Related projects

Related capabilities

More articles