DevConf.US 2026

Demystifying Quantization: Accelerating Open-Source LLM Inference with Red Hat AI
2026-09-24 –, Hewitt Boardroom (Capacity 35)

Quantization and model compression have become cornerstone techniques for accelerating large language model (LLM) inference, delivering impressive efficiency gains while reducing compute costs. Yet, for many practitioners, the field’s rapid evolution, breadth of algorithms, and technical depth can be intimidating. This session offers an accessible, first-principles introduction to quantization, helping attendees understand both the why and how behind it. In this session, we will:

  • Break down the fundamental concepts of quantization and explain the most widely used quantization formats.
  • Understand how quantization impacts model efficiency, accuracy, and performance trade-offs across different LLM architectures.
  • Explore advanced algorithms and techniques for squeezing the most value out of model performance while protecting model behavior
  • Demonstrate how to use open-source tools such as LLM Compressor and vLLM to optimally serve models as performantly as possible

By the end of this session, you’ll understand how to select and tune quantization strategies to achieve optimal performance and accuracy recovery for your own deployments. You’ll also see how industry leaders—such as the Meta AI team behind Llama 4—use the same Red Hat technologies to efficiently optimize large mixture-of-experts (MoE) models.