Quantization
This page explains the weight formats shown in Pipette and how quantization affects quality, speed, and memory.
Quantization stores model weights at lower precision. It usually reduces model size and can improve runtime performance, with a possible loss in quality. The effect depends on the model, weight format, runtime, hardware, and workload.
Lower bit width does not guarantee higher speed. The runtime needs optimized kernels for the exact quantization format on the target CPU, GPU, or accelerator.
Formats in Pipette
Formats below are ordered from highest to lowest bit width. The table is derived from the same quantization registry used by the dashboard.
| Quant | Bits | Notes |
|---|---|---|
fp16 | 16 | Half precision. Full quality baseline. |
bf16 | 16 | Brain Floating Point 16. Same bit width as fp16, with a wider dynamic range but less mantissa precision. Often preferred for inference. |
q8_0 | 8 | GGUF Q8_0. Signed 8-bit weights with one symmetric scale per 32-weight block. |
8bit | 8 | MLX 8-bit. Comparable to fp16 accuracy at ~half the size; mlx-lm format. |
q6_k | 6 | k-quant 6-bit. Close to q8_0 quality; mid-tier size. |
q5_k_m | 5 | k-quant medium. |
q4_k_m | 4 | k-quant medium. Better quality than q4_0 thanks to mixed precision within blocks. |
4bit | 4 | MLX 4-bit. mlx-lm equivalent of GGUF q4 variants; group-wise quantization. |
q4_0 | 4 | Basic 4-bit GGUF format: one scale per 32-weight block. Compact baseline, usually lower quality than Q4_K_M. |
q2_g64 | 2 | Sub-4-bit. Included because a launch model is published at it, not as a ladder rung. |
iq1_m | 1.75 | Sub-4-bit. Included because a launch model is published at it, not as a ladder rung. |
q1_0 | 1 | Sub-4-bit. Included because a launch model is published at it, not as a ladder rung. |
Full-precision references can use FP16 or BF16, depending on the native weights published for a model. These are both 16-bit formats, but they allocate their bits differently. Treat each as the full-precision reference for its model, not as one identical format shared by every model.
Choose a deployment point
The highest-precision model is not always the best deployment choice. A smaller quantization may be preferable if it fits the target device, responds faster, and retains enough quality for the task.
Use Compare Pipette Results to compare quality, speed, and memory together. Use Coverage and Selection to understand which weight sources and quantizations are eligible for public charts.