BEGIN:VCALENDAR
VERSION:2.0
PRODID:-//pretalx//pretalx.devconf.info//devconf-us-2026//speaker//SXYZK8
BEGIN:VTIMEZONE
TZID:EST
BEGIN:STANDARD
DTSTART:20001029T030000
RRULE:FREQ=YEARLY;BYDAY=-1SU;BYMONTH=10;UNTIL=20061029T070000Z
TZNAME:EST
TZOFFSETFROM:-0400
TZOFFSETTO:-0500
END:STANDARD
BEGIN:STANDARD
DTSTART:20071104T030000
RRULE:FREQ=YEARLY;BYDAY=1SU;BYMONTH=11
TZNAME:EST
TZOFFSETFROM:-0400
TZOFFSETTO:-0500
END:STANDARD
BEGIN:DAYLIGHT
DTSTART:20000402T030000
RRULE:FREQ=YEARLY;BYDAY=1SU;BYMONTH=4;UNTIL=20060402T080000Z
TZNAME:EDT
TZOFFSETFROM:-0500
TZOFFSETTO:-0400
END:DAYLIGHT
BEGIN:DAYLIGHT
DTSTART:20070311T030000
RRULE:FREQ=YEARLY;BYDAY=2SU;BYMONTH=3
TZNAME:EDT
TZOFFSETFROM:-0500
TZOFFSETTO:-0400
END:DAYLIGHT
END:VTIMEZONE
BEGIN:VEVENT
UID:pretalx-devconf-us-2026-F9HAXH@pretalx.devconf.info
DTSTART;TZID=EST:20260924T154000
DTEND;TZID=EST:20260924T155500
DESCRIPTION:Quantization and model compression have become cornerstone tech
 niques for accelerating large language model (LLM) inference\, delivering 
 impressive efficiency gains while reducing compute costs. Yet\, for many p
 ractitioners\, the field’s rapid evolution\, breadth of algorithms\, and
  technical depth can be intimidating. This session offers an accessible\, 
 first-principles introduction to quantization\, helping attendees understa
 nd both the why and how behind it. In this session\, we will:\n\n- Break d
 own the fundamental concepts of quantization and explain the most widely u
 sed quantization formats.\n- Understand how quantization impacts model eff
 iciency\, accuracy\, and performance trade-offs across different LLM archi
 tectures.\n- Explore advanced algorithms and techniques for squeezing the 
 most value out of model performance while protecting model behavior\n- Dem
 onstrate how to use open-source tools such as LLM Compressor and vLLM to o
 ptimally serve models as performantly as possible\n\nBy the end of this se
 ssion\, you’ll understand how to select and tune quantization strategies
  to achieve optimal performance and accuracy recovery for your own deploym
 ents. You’ll also see how industry leaders—such as the Meta AI team be
 hind Llama 4—use the same Red Hat technologies to efficiently optimize l
 arge mixture-of-experts (MoE) models.
DTSTAMP:20260924T235818Z
LOCATION:Hewitt Boardroom (Capacity 35)
SUMMARY:Demystifying Quantization: Accelerating Open-Source LLM Inference w
 ith Red Hat AI - Kyle Sayers
URL:https://pretalx.devconf.info/devconf-us-2026/talk/F9HAXH/
END:VEVENT
END:VCALENDAR
