LLM Quantization: Bits, Memory and Error

黎 浩然/ 11 10 月, 2026/ 大语言模型/LARGELANGUAGEMODEL/LLM, 机器学习/MACHINELEARNING, 研究生/POSTGRADUATE, 计算机/COMPUTER/ 0 comments

Quantizing LLM weights to eight or four bits can reduce weight storage, but a “4-bit model” does not imply four-bit computation throughout inference. Separate weights, computation and runtime memory before estimating savings or interpreting errors.

Contents
  1. Count raw weights, not total VRAM
  2. How integer codes reconstruct approximate values
  3. Clipping deserves its own check
  4. Storage bits and compute precision differ
  5. Compare the same tasks before deployment
  6. Sources

Count raw weights, not total VRAM

Suppose a model has exactly seven billion parameters, all tightly stored at one bit width with no additional information. Raw storage is parameter count × bits / 8 bytes: 14 GB at 16 bits, 7 GB at eight bits and 3.5 GB at four bits. These are decimal GB, with 1 GB=10⁹ bytes.

16-bit: 14 GB

8-bit: 7 GB

4-bit: 3.5 GB

Original raw-weight storage comparison for an invented seven-billion-parameter count. Decimal GB; excludes metadata and runtime memory.

Real quantized files also store information such as scales, and some layers may retain higher precision. A running service additionally needs KV cache, activations, workspaces and framework allocations. The 3.5 GB figure is raw bit storage under the stated assumptions, not proof that a machine can run the model.

Use the earlier KV-cache memory estimate to track weight and context budgets separately. Changing sequence length or concurrency can change peak runtime memory even with identical weights.

How integer codes reconstruct approximate values

A simple uniform affine mapping is q = clamp(round(x/s) + z), with reconstruction x̂ = s(q−z). The scale s is positive and z is an integer zero point. PyTorch’s official quantized-tensor explanation describes this scale-and-zero-point representation. It is one family of quantization, not a description of every low-bit LLM format.

Take s=0.02, z=0 and integer codes from −128 to 127. Adjacent reconstructed values differ by 0.02. Without clipping, nearest rounding gives a scalar absolute error of at most 0.01. This bound says nothing about model output quality and excludes clipped values.

Input x Code q Reconstruction
−1.16 −58 −1.16
−0.213 −11 −0.22
0.371 19 0.38
1.891 95 1.90

Clipping deserves its own check

The largest reconstruction in this mapping is 127×0.02=2.54. An input of 10 is clipped to code 127, giving an error of 7.46 rather than at most 0.01. Range selection trades coverage of extreme values against resolution for common values.

This original Python example was executed, checking four unclipped errors and the clipped result. It calls neither PyTorch nor bitsandbytes nor a real LLM. Python rounds exact halfway cases to even; the tested values avoid that boundary.

import math

def quantize(x, scale=0.02):
    if not math.isfinite(x) or not math.isfinite(scale) or scale <= 0:
        raise ValueError("finite x and a positive finite scale required")
    q = max(-128, min(127, round(x / scale)))
    return q, q * scale

values = [-1.16, -0.213, 0.371, 1.891]
for x in values:
    q, restored = quantize(x)
    assert abs(x - restored) <= 0.01 + 1e-12
    print(x, q, round(restored, 3))
q, restored = quantize(10.0)
assert q == 127 and math.isclose(restored, 2.54)
assert abs(10.0 - restored) > 0.01
print("clipped:", q, restored)

Including a large outlier in a global range can increase the step size, making approximations coarser elsewhere. A small constructed tensor can demonstrate that mechanism, but cannot establish whole-model quality. Grouping, channel-specific scaling or selected higher-precision computation require validation for the actual method.

Storage bits and compute precision differ

The Transformers v4.50.0 documentation provides a separate compute-dtype setting for four-bit loading and describes higher-precision handling of sensitive components in its eight-bit method. Low-bit storage therefore does not imply one uniform low-bit integer format for all arithmetic.

Reconstructing a floating-point tensor also does not recover the original values. A code q represents an approximate interval; discarded detail has not returned. Speed depends on kernels, hardware, shapes and conversion costs. Halving bit width is not a promise to halve latency.

Compare the same tasks before deployment

Keep the model and input/output length distributions fixed. Record completion rate, task quality, peak memory, time to first token and generation intervals before and after quantization. Preserve specific failures for numerical, coding and strict-format tasks: similar averages can hide more errors in one category.

When training adapters, distinguish base-weight storage from LoRA trainable parameters. This article checks bit-width arithmetic and uniform-quantization error. It reproduces no paper benchmark and selects no universally optimal bit width.

中文版

Sources

PyTorch: Introducing Quantized Tensor; Transformers v4.50.0: bitsandbytes.

Publication note: actually published as a catch-up on October 11, 2026 (Beijing time), retaining the planned 09:00 article date. Versioned documentation is used for historical feature descriptions, not as current installation guidance.

Support

If this article helped you, you can support this site.

WeChat support QR code; click to enlarge
WeChat
Alipay support QR code; click to enlarge
Alipay
Buy Me a Coffee; support this site
Buy Me a Coffee

Click a QR code to enlarge. More options: support page。

Share this Post

Leave a Comment

您的邮箱地址不会被公开。 必填项已用 * 标注

*
*