Serving 2 Distributed Training and Serving Architecture for LLMs: What Engineers Need to Know Dec 13, 2025 Inference Optimization for LLMs: Latency, Throughput, Quantization, and Serving Trade-Offs Dec 11, 2025