BEGIN:VCALENDAR
VERSION:2.0
PRODID:-//pretalx//pretalx.devconf.info//devconf-us-2026//talk//9SMF9X
BEGIN:VTIMEZONE
TZID:EST
BEGIN:STANDARD
DTSTART:20001029T030000
RRULE:FREQ=YEARLY;BYDAY=-1SU;BYMONTH=10;UNTIL=20061029T070000Z
TZNAME:EST
TZOFFSETFROM:-0400
TZOFFSETTO:-0500
END:STANDARD
BEGIN:STANDARD
DTSTART:20071104T030000
RRULE:FREQ=YEARLY;BYDAY=1SU;BYMONTH=11
TZNAME:EST
TZOFFSETFROM:-0400
TZOFFSETTO:-0500
END:STANDARD
BEGIN:DAYLIGHT
DTSTART:20000402T030000
RRULE:FREQ=YEARLY;BYDAY=1SU;BYMONTH=4;UNTIL=20060402T080000Z
TZNAME:EDT
TZOFFSETFROM:-0500
TZOFFSETTO:-0400
END:DAYLIGHT
BEGIN:DAYLIGHT
DTSTART:20070311T030000
RRULE:FREQ=YEARLY;BYDAY=2SU;BYMONTH=3
TZNAME:EDT
TZOFFSETFROM:-0500
TZOFFSETTO:-0400
END:DAYLIGHT
END:VTIMEZONE
BEGIN:VEVENT
UID:pretalx-devconf-us-2026-9SMF9X@pretalx.devconf.info
DTSTART;TZID=EST:20260925T104000
DTEND;TZID=EST:20260925T105500
DESCRIPTION:**Are you deploying Large Language Models and struggling to squ
 eeze every ounce of performance out of your costly GPU infrastructure?** W
 hile engines like vLLM are industry standards\, using their default config
 urations for your unique complex workload shape \\- involving long prefill
 \, extreme sequence lengths\, or high concurrency \\-  can often leave per
 formance on the table. Traditionally\, tuning these parameters meant relyi
 ng on manual tuning or black-box Bayesian Optimization. But what if the tu
 ning could *reason* about why a configuration failed?\n\nThis session intr
 oduces the next evolution of **llm-tuna**\, an open-source framework for a
 utomated LLM serving optimization. Building on our research published in t
 he ACM ([link](https://dl.acm.org/doi/10.1145/3774904.3792953))\, we intro
 duce an **agentic feedback loop** where an AI agent manages a persistent P
 erformance Knowledge Base. Instead of just navigating a parameter space\, 
 the agent proactively:\n\n1. **Reads and Interprets:** Analyzes vLLM logs 
 and engine metrics to identify memory pressure or compute bottlenecks.  \n
 2. **Profiles and Diagnoses:** Triggers low-level **PyTorch profiling** an
 d kernel-level analysis to understand *why* certain batch sizes or CUDA gr
 aph settings underperform.   \n3. **Learns and Adapts:** Grows its knowled
 ge base with every experiment\, allowing it to bypass inefficient search p
 aths and suggest optimizations that pure statistical models might miss and
  accelerate the search for newer users of the framework.\n\nWe will showca
 se the architecture—leveraging a kubernetes cluster for distributed\, is
 olated trials—and demonstrate how this "Expert-in-the-Loop" approach out
 performs standard baselines. This study will show how the progression impr
 oves over traditional statistical approaches. Attendees will see real-worl
 d case studies on models like **Qwen2.5** and the massive **DeepSeek-R1-67
 1B** on H100/H200 clusters.\n\n### **Key Takeaways:**\n\n* **The Agentic E
 dge:** How a reasoning agent identifies bottlenecks that Bayesian Optimiza
 tion ignores.  \n* **Closing the Loop:** Integrating PyTorch profiling and
  log analysis into the automated tuning cycle.  \n* **Building a Performan
 ce Moat:** Leveraging a persistent\, research-backed knowledge base to acc
 elerate deployments across diverse GPU architectures (H100/H200).  \n* **M
 ulti-Objective Optimization:** Balancing the trade-offs between prefill la
 tency\, decode throughput\, and VRAM efficiency.
DTSTAMP:20260727T175018Z
LOCATION:Ladd Room (Capacity 170)
SUMMARY:LLM-Tuna 2.0: Agentic Profiling and Knowledge-Driven Optimization f
 or vLLM - Thameem Abbas Ibrahim Bathusha\, Aanya Sharma
URL:https://pretalx.devconf.info/devconf-us-2026/talk/9SMF9X/
END:VEVENT
END:VCALENDAR
