DevConf.US 2026

Hands-On llm-d: Building High-Performance AI Inference at Scale
2026-09-24 –, 107 (Capacity 20)

Llm-d gives you the ability to deliver generative AI capability at scale at a higher density and performance given the same amount of hardware, if supported by the architecture.

Technical attendees will learn not only how to perform these actions, but how to demonstrate their value through repeatable steps that show measurable improvements in throughput and latency. Specifically, attendees will get hands on experience with:
- Intelligent Inference Scheduling and its benefits achieving your SLOs on a limited hardware
- Designing, running, and understanding benchmarks with GuideLLM

Attendees will also gain an understanding of Prefill/Decode disaggregation for scaling phases of inference independently and Wide Expert Parallelism for MoE models like DeepSeek R1, understanding how to serve the largest models at scale.