Concept
Model Quantization
Reducing weight precision from FP16 or FP32 toward INT8 or INT4 shrinks VRAM footprints and speeds matmuls. Accuracy shifts per task, so benchmark perplexity or task metrics before production. Serving stacks expose flags to pick precision per deployment tier.
1 documentation page cover this concept. Editorial glossary