BEGIN:VCALENDAR
VERSION:2.0
PRODID:-//pretalx//pretalx.devconf.info//devconf-us-2026//speaker//Y77XKM
BEGIN:VTIMEZONE
TZID:EST
BEGIN:STANDARD
DTSTART:20001029T030000
RRULE:FREQ=YEARLY;BYDAY=-1SU;BYMONTH=10;UNTIL=20061029T070000Z
TZNAME:EST
TZOFFSETFROM:-0400
TZOFFSETTO:-0500
END:STANDARD
BEGIN:STANDARD
DTSTART:20071104T030000
RRULE:FREQ=YEARLY;BYDAY=1SU;BYMONTH=11
TZNAME:EST
TZOFFSETFROM:-0400
TZOFFSETTO:-0500
END:STANDARD
BEGIN:DAYLIGHT
DTSTART:20000402T030000
RRULE:FREQ=YEARLY;BYDAY=1SU;BYMONTH=4;UNTIL=20060402T080000Z
TZNAME:EDT
TZOFFSETFROM:-0500
TZOFFSETTO:-0400
END:DAYLIGHT
BEGIN:DAYLIGHT
DTSTART:20070311T030000
RRULE:FREQ=YEARLY;BYDAY=2SU;BYMONTH=3
TZNAME:EDT
TZOFFSETFROM:-0500
TZOFFSETTO:-0400
END:DAYLIGHT
END:VTIMEZONE
BEGIN:VEVENT
UID:pretalx-devconf-us-2026-JQK7PL@pretalx.devconf.info
DTSTART;TZID=EST:20260924T142000
DTEND;TZID=EST:20260924T145500
DESCRIPTION:Multimodal large language models introduce a class of inference
  problem that standard vLLM profiling playbooks were not designed for. Whi
 le LLM inference is well-understood\, Multimodal MoE models introduce a ch
 aotic new variable: image resolution. In a real-world product catalog\, 20
 x variance in image size disrupts standard batch formation\, drives preemp
 tion spikes.\nThis talk presents the profiling methodology and optimizatio
 n path we followed to achieve the top result among all B200 submissions on
  the MLPerf Inference benchmark for Qwen3-VL-235B-A22B-Instruct\, a 235-bi
 llion-parameter mixture-of-experts vision-language model with 22 billion a
 ctive parameters. The workload is the MLPerf Shopify product catalog: 48\,
 289 real e-commerce images paired with text\, classified into structured J
 SON output. We submitted results on both 8xB200 and 8xH200 nodes running R
 HEL with CentML-optimized vLLM.\n\nResults\nConfiguration\nThroughput\nCon
 text\n8×B200 — Server\n67.86 samples/sec\n #1 among all B200 submission
 s\; 50% faster than top B300 submission\n8×B200 — Offline\n79.04 sample
 s/sec\n #1 B200 result\; on par with top B300 submission\n8×H200 — Offl
 ine\n18.22 samples/sec\n Only H200 submission for this model\n8×H200 — 
 Server\n11.05 samples/sec\n Only H200 submission for this model\n\nWe will
  present the profiling work that drove these results and how we identified
  the ViT encoder as the primary throughput constraint\, why Shortest Job F
 irst scheduling outperformed continuous batching defaults on variable-imag
 e workloads\, how FP8 multimodal attention was validated without compromis
 ing benchmark accuracy\, and what the FlashInfer MoE kernel traces reveale
 d about active-parameter utilization under the MLPerf traffic pattern.
DTSTAMP:20260727T165258Z
LOCATION:Ladd Room (Capacity 170)
SUMMARY:Profiling and Optimizing vLLM for Large Multimodal MoE Inference - 
 HARIKA POTHINA\, Naveen Miriyalu
URL:https://pretalx.devconf.info/devconf-us-2026/talk/JQK7PL/
END:VEVENT
END:VCALENDAR
