DevConf.US 2026

The Hardware Hunger Games: Why Your AI Training Depends on Infrastructure SLOs
2026-09-25 , 101 (Capacity 48)

Managing Scarce ML Accelerator Resources

The demand for TPUs and GPUs to power Large Language Models is exploding, far outpacing the months-long lead times for new hardware. This creates intense competition for scarce resources. The critical question for infrastructure teams is: How do you guarantee reliable, predictable, and fair access to chips that take months to provision? Simply measuring hardware uptime is insufficient when users care about getting their ML workloads completed.

This session will dissect the complexities of managing platform-agnostic large-scale ML accelerator fleets, independent of the specific hardware type (TPUs, GPUs, ASICs) within open-source platforms like Kubernetes or for running whale customer LLM workloads like Gemini. We will cover:

  • Holistic SLO Framework: Defining, measuring, and defending meaningful Capacity SLOs (e.g., available chip hours, scheduling latency) and Workload SLOs (e.g., Goodput, training throughput).
  • Multi-Tenant Scheduling: Strategies for fair share, priority-based queuing, and minimizing head-of-line blocking.
  • Hardware Fragmentation: Identifying and mitigating internal and external fragmentation that reduces usable capacity.
  • Preemption Logic: Designing and implementing preemption mechanisms to enforce SLOs and priorities, including the impact on workload Goodput.
  • Quota Management & Demand Shaping: Best practices for assigning quotas and influencing user behavior to smooth out peaks.
  • Handling Burst Requests: Strategies for accommodating episodic high-demand training cycles without compromising base SLOs.

We will explore the Site Reliability Engineering practices essential for robust accelerator fleet management:

  • Establishing Realistic SLOs: How to set achievable targets for capacity and workload performance.
  • Error Budgets for Accelerators: Implementing effective error budgets tailored to capacity, demand fluctuations, and workload characteristics.
  • Monitoring & Alerting: Crucial metrics and strategies to track SLO performance, detect saturation, identify fragmentation, and diagnose workload-specific issues.
  • Communication: Transparently communicating resource constraints, SLOs, and the impact of burst requests to users.
  • Balancing Efficiency & Fairness: Techniques for optimizing fleet utilization in multi-tenant environments with diverse requirements.

Attendees will leave with a robust, platform-agnostic framework for implementing holistic Capacity and Workload SLOs for high-demand ML accelerators. You'll gain actionable SRE principles and techniques to enhance reliability and techniques to enhance t


What level of experience should the audience have to best understand your session?: Beginner - no experience needed

Sania Alex is a Senior Software Engineer at Google, leading Optical Circuit Switch operations for optimizing workload scheduling and capacity on TPUs (Tensor Processing Units). Her team is responsible for defending infrastructure SLOs for unified, efficient, and reliable ML production infrastructure (TPUs and GPUs) across Alphabet. She is a seasoned Software Engineer with 8 years of expertise in software development, cloud computing, and machine learning. Her impressive career includes delivering key projects for Amazon Web Services' Simple Storage Service (S3). Sania specializes at optimizing ML resource usage and capacity sharing across Google's powerful ML hardware, which underpins core Google products like Gemini, a large language model powering Google Search and other products. When not working to enhance the reliability and efficiency of ML hardware, Sania enjoys spending time with her family and immersing herself in a good book.

Swetha Vijayaraghavan is a software engineer with 18+ years of experience. She has worked on building sofware applications for various domains including eCommerce, Finance and Infrastuecture. Currently she works as an Infrastructure SRE building, tools, dashboards and automation for the reliability ML platforms within Google.