BEGIN:VCALENDAR
VERSION:2.0
PRODID:-//pretalx//pretalx.devconf.info//devconf-us-2026//talk//DE7B7X
BEGIN:VTIMEZONE
TZID:EST
BEGIN:STANDARD
DTSTART:20001029T030000
RRULE:FREQ=YEARLY;BYDAY=-1SU;BYMONTH=10;UNTIL=20061029T070000Z
TZNAME:EST
TZOFFSETFROM:-0400
TZOFFSETTO:-0500
END:STANDARD
BEGIN:STANDARD
DTSTART:20071104T030000
RRULE:FREQ=YEARLY;BYDAY=1SU;BYMONTH=11
TZNAME:EST
TZOFFSETFROM:-0400
TZOFFSETTO:-0500
END:STANDARD
BEGIN:DAYLIGHT
DTSTART:20000402T030000
RRULE:FREQ=YEARLY;BYDAY=1SU;BYMONTH=4;UNTIL=20060402T080000Z
TZNAME:EDT
TZOFFSETFROM:-0500
TZOFFSETTO:-0400
END:DAYLIGHT
BEGIN:DAYLIGHT
DTSTART:20070311T030000
RRULE:FREQ=YEARLY;BYDAY=2SU;BYMONTH=3
TZNAME:EDT
TZOFFSETFROM:-0500
TZOFFSETTO:-0400
END:DAYLIGHT
END:VTIMEZONE
BEGIN:VEVENT
UID:pretalx-devconf-us-2026-DE7B7X@pretalx.devconf.info
DTSTART;TZID=EST:20260924T130000
DTEND;TZID=EST:20260924T133500
DESCRIPTION:Migrating legacy documentation often reveals a critical bottlen
 eck for enterprise operations: trapped\, unstructured data. Red Hat’s Pr
 ocurement team recently faced this challenge when migrating over 32\,000 l
 egacy vendor contracts to a new Contract Lifecycle Manager (CLM). The new 
 system required specific metadata - such as creation dates\, expiration da
 tes\, and signatories - extracted from a complex mix of generated and scan
 ned PDFs. These documents spanned multiple languages and featured non-stan
 dard layouts and tabular data. Completing this extraction manually was pro
 jected to take over a year\, consuming approximately 14\,000 hours and $70
 0\,000 in operational costs.\n\nThis session details how Red Hat engineere
 d a reusable\, AI-driven automation pipeline to solve this challenge in a 
 fraction of the time. Hosted on Red Hat OpenShift AI\, the solution utiliz
 es Docling to convert highly unstructured PDFs into Markdown. From there\,
  the text is processed through a Qwen2.5 32B LLM to intelligently extract 
 and format the required metadata into structured JSON.\n\nBy transitioning
  to an automated\, unstructured-to-structured ETL approach\, the team comp
 leted the migration in just 3 months. The project required only 1\,700 hou
 rs of development and review\, ultimately saving 75% in FTE hours and achi
 eving $600\,000 in cost savings.
DTSTAMP:20260727T175219Z
LOCATION:Ladd Room (Capacity 170)
SUMMARY:Unstructured-to-Structured ETL: Automating Contract Metadata Extrac
 tion with Docling and GenAI - Taylor Agarwal
URL:https://pretalx.devconf.info/devconf-us-2026/talk/DE7B7X/
END:VEVENT
END:VCALENDAR
