BEGIN:VCALENDAR
VERSION:2.0
PRODID:-//pretalx//pretalx.devconf.info//devconf-us-2026//talk//MXQC9E
BEGIN:VTIMEZONE
TZID:EST
BEGIN:STANDARD
DTSTART:20001029T030000
RRULE:FREQ=YEARLY;BYDAY=-1SU;BYMONTH=10;UNTIL=20061029T070000Z
TZNAME:EST
TZOFFSETFROM:-0400
TZOFFSETTO:-0500
END:STANDARD
BEGIN:STANDARD
DTSTART:20071104T030000
RRULE:FREQ=YEARLY;BYDAY=1SU;BYMONTH=11
TZNAME:EST
TZOFFSETFROM:-0400
TZOFFSETTO:-0500
END:STANDARD
BEGIN:DAYLIGHT
DTSTART:20000402T030000
RRULE:FREQ=YEARLY;BYDAY=1SU;BYMONTH=4;UNTIL=20060402T080000Z
TZNAME:EDT
TZOFFSETFROM:-0500
TZOFFSETTO:-0400
END:DAYLIGHT
BEGIN:DAYLIGHT
DTSTART:20070311T030000
RRULE:FREQ=YEARLY;BYDAY=2SU;BYMONTH=3
TZNAME:EDT
TZOFFSETFROM:-0500
TZOFFSETTO:-0400
END:DAYLIGHT
END:VTIMEZONE
BEGIN:VEVENT
UID:pretalx-devconf-us-2026-MXQC9E@pretalx.devconf.info
DTSTART;TZID=EST:20260925T090000
DTEND;TZID=EST:20260925T093500
DESCRIPTION:## Managing Scarce ML Accelerator Resources\n\nThe demand for T
 PUs and GPUs to power Large Language Models is exploding\, far outpacing t
 he months-long lead times for new hardware. This creates intense competiti
 on for scarce resources. The critical question for infrastructure teams is
 : How do you guarantee reliable\, predictable\, and fair access to chips t
 hat take months to provision? Simply measuring hardware uptime is insuffic
 ient when users care about getting their ML workloads completed.\n\nThis s
 ession will dissect the complexities of managing platform-agnostic large-s
 cale ML accelerator fleets\, independent of the specific hardware type (TP
 Us\, GPUs\, ASICs) within open-source platforms like Kubernetes or for run
 ning whale customer LLM workloads like Gemini. We will cover:\n\n*   **Hol
 istic SLO Framework:** Defining\, measuring\, and defending meaningful Cap
 acity SLOs (e.g.\, available chip hours\, scheduling latency) and Workload
  SLOs (e.g.\, Goodput\, training throughput).\n*   **Multi-Tenant Scheduli
 ng:** Strategies for fair share\, priority-based queuing\, and minimizing 
 head-of-line blocking.\n*   **Hardware Fragmentation:** Identifying and mi
 tigating internal and external fragmentation that reduces usable capacity.
 \n*   **Preemption Logic:** Designing and implementing preemption mechanis
 ms to enforce SLOs and priorities\, including the impact on workload Goodp
 ut.\n*   **Quota Management & Demand Shaping:** Best practices for assigni
 ng quotas and influencing user behavior to smooth out peaks.\n*   **Handli
 ng Burst Requests:** Strategies for accommodating episodic high-demand tra
 ining cycles without compromising base SLOs.\n\nWe will explore the Site R
 eliability Engineering practices essential for robust accelerator fleet ma
 nagement:\n\n*   **Establishing Realistic SLOs:** How to set achievable ta
 rgets for capacity and workload performance.\n*   **Error Budgets for Acce
 lerators:** Implementing effective error budgets tailored to capacity\, de
 mand fluctuations\, and workload characteristics.\n*   **Monitoring & Aler
 ting:** Crucial metrics and strategies to track SLO performance\, detect s
 aturation\, identify fragmentation\, and diagnose workload-specific issues
 .\n*   **Communication:** Transparently communicating resource constraints
 \, SLOs\, and the impact of burst requests to users.\n*   **Balancing Effi
 ciency & Fairness:** Techniques for optimizing fleet utilization in multi-
 tenant environments with diverse requirements.\n\nAttendees will leave wit
 h a robust\, platform-agnostic framework for implementing holistic Capacit
 y and Workload SLOs for high-demand ML accelerators. You'll gain actionabl
 e SRE principles and techniques to enhance reliability and techniques to e
 nhance t
DTSTAMP:20260727T174807Z
LOCATION:101 (Capacity 48)
SUMMARY:The Hardware Hunger Games: Why Your AI Training Depends on Infrastr
 ucture SLOs - Sania Alex\, Swetha Vijayaraghavan
URL:https://pretalx.devconf.info/devconf-us-2026/talk/MXQC9E/
END:VEVENT
END:VCALENDAR
