BEGIN:VCALENDAR
VERSION:2.0
PRODID:-//pretalx//pretalx.devconf.info//devconf-us-2026//speaker//GED7L8
BEGIN:VTIMEZONE
TZID:EST
BEGIN:STANDARD
DTSTART:20001029T030000
RRULE:FREQ=YEARLY;BYDAY=-1SU;BYMONTH=10;UNTIL=20061029T070000Z
TZNAME:EST
TZOFFSETFROM:-0400
TZOFFSETTO:-0500
END:STANDARD
BEGIN:STANDARD
DTSTART:20071104T030000
RRULE:FREQ=YEARLY;BYDAY=1SU;BYMONTH=11
TZNAME:EST
TZOFFSETFROM:-0400
TZOFFSETTO:-0500
END:STANDARD
BEGIN:DAYLIGHT
DTSTART:20000402T030000
RRULE:FREQ=YEARLY;BYDAY=1SU;BYMONTH=4;UNTIL=20060402T080000Z
TZNAME:EDT
TZOFFSETFROM:-0500
TZOFFSETTO:-0400
END:DAYLIGHT
BEGIN:DAYLIGHT
DTSTART:20070311T030000
RRULE:FREQ=YEARLY;BYDAY=2SU;BYMONTH=3
TZNAME:EDT
TZOFFSETFROM:-0500
TZOFFSETTO:-0400
END:DAYLIGHT
END:VTIMEZONE
BEGIN:VEVENT
UID:pretalx-devconf-us-2026-VRR9M8@pretalx.devconf.info
DTSTART;TZID=EST:20260925T110000
DTEND;TZID=EST:20260925T113500
DESCRIPTION:Most agent security testing uses models that refuse malicious p
 rompts\, so infrastructure controls never get exercised. We tested which O
 penShift defenses actually hold when the model says yes to everything.\nWe
  deployed OpenClaw on Red Hat OpenShift (ROKS) with an abliterated Qwen3.5
 -35B-A3B served via vLLM: 100% attack cooperation\, zero refusals. TrojAI 
 calls this "crash test dummy" methodology. The model cooperates with every
  attack\, forcing each infrastructure layer to prove itself independently.
  We wrote 15 custom garak probes across six attack categories (credential 
 exfiltration\, persistence poisoning\, sandbox escape\, tool abuse\, Kuber
 netes API escalation\, guardrails bypass) and ran 91 adversarial prompts a
 gainst three progressive hardening tiers: bare agent\, sandbox with Networ
 kPolicy\, and sandbox with a prompt injection classifier.\nThe results sep
 arate load-bearing controls from cosmetic ones. Credential isolation (rout
 ing tool execution to a separate pod) dropped exfiltration from 67% to 0%.
  NetworkPolicy blocked K8s API abuse\, but only after we discovered that i
 pBlock rules can't target ClusterIPs because kube-proxy DNAT translates be
 fore policy evaluation. DNS tunneling worked until we restricted egress to
  cluster DNS pods by label. A prompt injection classifier caught encoding-
 based attacks (56% to 0%) but missed every tool-use attack phrased as a le
 gitimate instruction. Persistence poisoning\, where the agent writes attac
 ker content into its own memory\, passed all three tiers because it looks 
 identical to normal agent behavior. No tested control addresses it.\n\nAtt
 endees leave with the three-tier testing framework\, garak probe configura
 tions\, and a checklist for validating agent isolation on Kubernetes. All 
 probes\, scan scripts\, and deployment manifests are open source.
DTSTAMP:20260727T165400Z
LOCATION:106 (Capacity 45)
SUMMARY:Red Teaming AI Agents on OpenShift: What Actually Stops the Attacks
  - Roy Belio
URL:https://pretalx.devconf.info/devconf-us-2026/talk/VRR9M8/
END:VEVENT
END:VCALENDAR
