BEGIN:VCALENDAR
VERSION:2.0
PRODID:-//pretalx//pretalx.devconf.info//devconf-us-2026//speaker//LBNCTV
BEGIN:VTIMEZONE
TZID:EST
BEGIN:STANDARD
DTSTART:20001029T030000
RRULE:FREQ=YEARLY;BYDAY=-1SU;BYMONTH=10;UNTIL=20061029T070000Z
TZNAME:EST
TZOFFSETFROM:-0400
TZOFFSETTO:-0500
END:STANDARD
BEGIN:STANDARD
DTSTART:20071104T030000
RRULE:FREQ=YEARLY;BYDAY=1SU;BYMONTH=11
TZNAME:EST
TZOFFSETFROM:-0400
TZOFFSETTO:-0500
END:STANDARD
BEGIN:DAYLIGHT
DTSTART:20000402T030000
RRULE:FREQ=YEARLY;BYDAY=1SU;BYMONTH=4;UNTIL=20060402T080000Z
TZNAME:EDT
TZOFFSETFROM:-0500
TZOFFSETTO:-0400
END:DAYLIGHT
BEGIN:DAYLIGHT
DTSTART:20070311T030000
RRULE:FREQ=YEARLY;BYDAY=2SU;BYMONTH=3
TZNAME:EDT
TZOFFSETFROM:-0500
TZOFFSETTO:-0400
END:DAYLIGHT
END:VTIMEZONE
BEGIN:VEVENT
UID:pretalx-devconf-us-2026-FQWK8Y@pretalx.devconf.info
DTSTART;TZID=EST:20260924T112000
DTEND;TZID=EST:20260924T115500
DESCRIPTION:Your alert fires at 3am. An SRE investigates\, finds the root c
 ause\, applies a known fix\, and goes back to sleep. Next week\, the same 
 alert fires again. Same investigation. Same fix. Same person woken up.\n\n
 **What if the cluster could do this itself** — investigate the alert\, r
 eason about the root cause\, select the right remediation\, execute it\, a
 nd produce an audit trail that proves it made the right call?\n\nThis talk
  explores the engineering challenges of building closed-loop Kubernetes se
 lf-healing with Agents as the reasoning layer. We'll dig into the hard pro
 blems:\n\n- **Trust boundaries**: How do you let an Agent investigate a pr
 oduction cluster without exposing secrets? How do you scope RBAC so the AI
  agent can read what it needs and nothing more?\n- **Auditability**: How d
 o you make "AI decided to scale your HPA" auditable enough for SOC2? What 
 does an immutable\, tamper-evident audit trail look like for autonomous re
 mediation?\n- **Governance**: How do you keep a human in the loop without 
 turning it into another page? When should the AI just fix it\, and when sh
 ould it ask?\n- **Fleet-scale credentials**: How do you architect ephemera
 l\, zero-trust credentials when the same pipeline needs to remediate acros
 s a fleet of clusters using ACM and AAP?\n\n### Live demo\n\nWe'll trigger
  a real alert on a Kind cluster and watch the full pipeline execute withou
 t human intervention:\n\n1. **Signal** → AlertManager fires\n2. **Invest
 igation** → an AI Agent queries the Kubernetes API\, reads logs\, descri
 bes pods\, builds a root cause analysis\n3. **Decision** →  an AI Agent 
 selects the right remediation workflow based on the RCA\n4. **Execution** 
 → K8s Jobs/Tekton/Ansible applies the fix\n5. **Audit** → Immutable ev
 ent records what happened and why\n\nNo slides-only theory. Bring your ske
 pticism.
DTSTAMP:20260727T165140Z
LOCATION:101 (Capacity 48)
SUMMARY:When Your Kubernetes Cluster Fixes Itself at 3am — And You Can Pr
 ove It Did the Right Thing - Jordi Gil\, Raghuram Banda
URL:https://pretalx.devconf.info/devconf-us-2026/talk/FQWK8Y/
END:VEVENT
END:VCALENDAR
