2026-09-24 –, Ladd Room (Capacity 170)
Performance engineers at Red Hat analyze vLLM inference benchmarks across Nvidia, AMD and TPU accelerators every release cycle. The data lives everywhere: benchmark CSVs, GPU metrics in Grafana, PyTorch profiler traces, vLLM source code on GitHub, and server logs with engine configuration details. Answering "why did throughput regress by 12% between vLLM-0.16.0 and vLLM-0.18.0 for model X?" used to mean hours of manual cross-referencing. We built an AI agent to unify it all and do it in minutes.
The AI Performance Agent is a LangGraph ReAct agent powered by Google Gemini, connected to a FastMCP server exposing 30 specialized tools for benchmark querying, kernel-level profiler analysis, source code diffing, GPU metrics, energy computation, cost analysis and more, all through natural language.
This talk covers the MCP architecture that makes the tools reusable across any client, the prompt engineering that teaches the model when to use which tool and when to stop and ask for clarification , the Langfuse v3 observability stack, and deploying the full system on OpenShift. Live demo included.
What attendees will take away:
- A reusable architecture pattern for building MCP-based AI agents that integrate with enterprise data sources
- Practical prompt engineering patterns for multi-tool agents that need to be rigorous, not just fluent
- A deployment blueprint for running LangGraph agents with full observability on OpenShift
- The open-source tool stack: LangGraph, FastMCP, Langfuse, Streamlit, all community projects
Performance & Scale Engineer at Red Hat, specializing in building and optimizing LLM inference at scale.
With a strong foundation in distributed systems, cloud computing, and MLOps, I bring expertise in designing, deploying, and optimizing large-scale, resilient infrastructure for modern AI and enterprise applications. Recognized for impactful open-source contributions and research publications at leading conferences, I’m passionate about pushing the boundaries of AI infrastructure to deliver high-performance, reliable systems.