BEGIN:VCALENDAR
VERSION:2.0
PRODID:-//pretalx//pretalx.devconf.info//devconf-us-2026//speaker//XVKQNH
BEGIN:VTIMEZONE
TZID:EST
BEGIN:STANDARD
DTSTART:20001029T030000
RRULE:FREQ=YEARLY;BYDAY=-1SU;BYMONTH=10;UNTIL=20061029T070000Z
TZNAME:EST
TZOFFSETFROM:-0400
TZOFFSETTO:-0500
END:STANDARD
BEGIN:STANDARD
DTSTART:20071104T030000
RRULE:FREQ=YEARLY;BYDAY=1SU;BYMONTH=11
TZNAME:EST
TZOFFSETFROM:-0400
TZOFFSETTO:-0500
END:STANDARD
BEGIN:DAYLIGHT
DTSTART:20000402T030000
RRULE:FREQ=YEARLY;BYDAY=1SU;BYMONTH=4;UNTIL=20060402T080000Z
TZNAME:EDT
TZOFFSETFROM:-0500
TZOFFSETTO:-0400
END:DAYLIGHT
BEGIN:DAYLIGHT
DTSTART:20070311T030000
RRULE:FREQ=YEARLY;BYDAY=2SU;BYMONTH=3
TZNAME:EDT
TZOFFSETFROM:-0500
TZOFFSETTO:-0400
END:DAYLIGHT
END:VTIMEZONE
BEGIN:VEVENT
UID:pretalx-devconf-us-2026-VKBKSS@pretalx.devconf.info
DTSTART;TZID=EST:20260925T110000
DTEND;TZID=EST:20260925T113500
DESCRIPTION:90% of GenAI systems fail in production because teams struggle 
 to evaluate and resolve failures. Observability tools like OpenTelemetry a
 nd MLflow log massive traces\, but hide critical reasoning failures. They 
 tell you where a request traveled\, but not why an LLM agent lost its way.
  This session moves beyond basic observability into the realm of failure m
 ode analysis for agentic systems. Using a real world AI system as a case s
 tudy\, we will dissect why traditional monitoring fails to catch silent ag
 ent failures like context window saturation and tool-calling loops.\n\nAtt
 endees will walk away with a practical toolkit for:\n1) Defining a Failure
  Taxonomy: Moving from vague bad outputs to structured categories like ret
 rieval gaps or reasoning drift\n2) Building a Tracing-to-Resolution Pipeli
 ne: Leveraging OpenTelemetry and MLflow traces to map anomalies directly t
 o code-level fixes\, eliminating the guesswork of manual prompt engineerin
 g\n3) Human-in-the-Loop Evaluations: Scaling quality control by using LLMs
  to surface high-risk anomalies for critical human review\n\nLearn to stop
  operating blind and build resilient\, production-grade AI systems that yo
 ur teams can actually trust.
DTSTAMP:20260727T165257Z
LOCATION:Ladd Room (Capacity 170)
SUMMARY:From Traces to Fixes: Building Resilient AgentOps with Failure Mode
  Analysis - Hema Veeradhi\, Surya Pathak
URL:https://pretalx.devconf.info/devconf-us-2026/talk/VKBKSS/
END:VEVENT
END:VCALENDAR
