BEGIN:VCALENDAR
VERSION:2.0
PRODID:-//pretalx//pretalx.devconf.info//devconf-us-2026//speaker//AQX3YB
BEGIN:VTIMEZONE
TZID:EST
BEGIN:STANDARD
DTSTART:20001029T030000
RRULE:FREQ=YEARLY;BYDAY=-1SU;BYMONTH=10;UNTIL=20061029T070000Z
TZNAME:EST
TZOFFSETFROM:-0400
TZOFFSETTO:-0500
END:STANDARD
BEGIN:STANDARD
DTSTART:20071104T030000
RRULE:FREQ=YEARLY;BYDAY=1SU;BYMONTH=11
TZNAME:EST
TZOFFSETFROM:-0400
TZOFFSETTO:-0500
END:STANDARD
BEGIN:DAYLIGHT
DTSTART:20000402T030000
RRULE:FREQ=YEARLY;BYDAY=1SU;BYMONTH=4;UNTIL=20060402T080000Z
TZNAME:EDT
TZOFFSETFROM:-0500
TZOFFSETTO:-0400
END:DAYLIGHT
BEGIN:DAYLIGHT
DTSTART:20070311T030000
RRULE:FREQ=YEARLY;BYDAY=2SU;BYMONTH=3
TZNAME:EDT
TZOFFSETFROM:-0500
TZOFFSETTO:-0400
END:DAYLIGHT
END:VTIMEZONE
BEGIN:VEVENT
UID:pretalx-devconf-us-2026-PQ8KNJ@pretalx.devconf.info
DTSTART;TZID=EST:20260925T153000
DTEND;TZID=EST:20260925T160500
DESCRIPTION:What if you could combine the speed of a small model with the q
 uality of a large one? Speculative decoding makes this possible by using a
  lightweight draft model to propose tokens that a larger verifier checks i
 n parallel\, delivering 2-5x inference speedups with mathematically identi
 cal quality output.\n\nThis talk introduces the vLLM project's Speculators
  library: an end-to-end framework for training\, packaging\, and deploying
  speculative decoding models. We will start with the intuition behind spec
 ulative decoding\, walk through the full Speculators training pipeline\, a
 nd finish with a live demonstration of serving a speculative decoding mode
 l to speed up inference.\n\nAlong the way\, we will dive into the engineer
 ing challenges we solved to make this work at scale. Extracting training d
 ata from large language models generates hundreds of megabytes per request
 \, which need to be transferred out of vLLM. Therefore we built a novel hi
 dden states extraction system that repurposes vLLM's KV Connector infrastr
 ucture to efficiently stream internal model representations to the trainin
 g processes\, all while preserving vLLM's tensor parallelism and paged att
 ention optimizations. We will share what we learned scaling this across mu
 lti-GPU setups and across speculative decoding algorithms like Eagle-3 and
  DFlash. \n\nAttendees will leave with a practical understanding of how sp
 eculative decoding works\, how to train their own draft models using Specu
 lators\, and how to deploy them for accelerated inference in vLLM.
DTSTAMP:20260727T165349Z
LOCATION:Ladd Room (Capacity 170)
SUMMARY:Speculators: Accelerating LLM Inference through Speculative Decodin
 g - Fynn Schmitt-Ulms
URL:https://pretalx.devconf.info/devconf-us-2026/talk/PQ8KNJ/
END:VEVENT
END:VCALENDAR
