Quantization 1 Inference Optimization for LLMs: Latency, Throughput, Quantization, and Serving Trade-Offs Dec 11, 2025