BEGIN:VCALENDAR
VERSION:2.0
PRODID:-//pretalx//pretalx.devconf.info//devconf-us-2026//talk//LVXBND
BEGIN:VTIMEZONE
TZID:EST
BEGIN:STANDARD
DTSTART:20001029T030000
RRULE:FREQ=YEARLY;BYDAY=-1SU;BYMONTH=10;UNTIL=20061029T070000Z
TZNAME:EST
TZOFFSETFROM:-0400
TZOFFSETTO:-0500
END:STANDARD
BEGIN:STANDARD
DTSTART:20071104T030000
RRULE:FREQ=YEARLY;BYDAY=1SU;BYMONTH=11
TZNAME:EST
TZOFFSETFROM:-0400
TZOFFSETTO:-0500
END:STANDARD
BEGIN:DAYLIGHT
DTSTART:20000402T030000
RRULE:FREQ=YEARLY;BYDAY=1SU;BYMONTH=4;UNTIL=20060402T080000Z
TZNAME:EDT
TZOFFSETFROM:-0500
TZOFFSETTO:-0400
END:DAYLIGHT
BEGIN:DAYLIGHT
DTSTART:20070311T030000
RRULE:FREQ=YEARLY;BYDAY=2SU;BYMONTH=3
TZNAME:EDT
TZOFFSETFROM:-0500
TZOFFSETTO:-0400
END:DAYLIGHT
END:VTIMEZONE
BEGIN:VEVENT
UID:pretalx-devconf-us-2026-LVXBND@pretalx.devconf.info
DTSTART;TZID=EST:20260925T102000
DTEND;TZID=EST:20260925T103500
DESCRIPTION:The explosive demand for Large Language Models has created a gl
 obal GPU bottleneck\, forcing platform engineers to seek alternative hardw
 are for production inference. Enter the vLLM engine on Google Cloud TPUs. 
 Traditionally optimized for GPUs\, vLLM now brings its industry-leading Pa
 gedAttention and continuous batching capabilities to the TPU ecosystem\, o
 ffering a high-throughput\, cost-effective alternative for serving models 
 like Llama-3.3-70B-Instruct.\n\nIn this session\, we dive into the archite
 cture of running vLLM on Google Kubernetes Engine (GKE). We will  demonstr
 ate the practicalities of deploying vLLM on TPU v6e slices and utilizing G
 KE for seamless orchestration.
DTSTAMP:20260727T175230Z
LOCATION:Ladd Room (Capacity 170)
SUMMARY:LLM inference using vLLM on TPU and GKE - Tahmid Muttaki
URL:https://pretalx.devconf.info/devconf-us-2026/talk/LVXBND/
END:VEVENT
END:VCALENDAR
