
Opening the Black Box: Tracking, Troubleshooting, and Auditing AI Agents in Production
We've all seen the dazzling demos of autonomous AI agents seamlessly executing complex workflows. But what happens when an agent breaks down in production? Moving AI out of the lab means dealing with the messy reality of troubleshooting systems that don't always behave predictably. It's no longer just about tracking traditional API response times; it's about auditing the "why" behind an agent's autonomous decisions.
In this session we'll skip the vendor lock-in and look at how to monitor agents using OpenTelemetry (OTEL) Gen AI Semantic Conventions. Through a live demonstration, we will show how to trace complex reasoning loops, capture asynchronous state changes, and build an audit trail that keeps both debugging engineers and compliance teams happy. You'll walk away with a pragmatic blueprint to track exactly what your agents are doing, why they're doing it, and how to prove it.
October 22 — Conference Day
Jerome is a highly experienced Cloud Architect, with more than 10 years experience in designing, building, and migrating workloads in the cloud. He has experience in both Azure and AWS cloud platforms enabling him to strategically evaluate the best option for a specific situation and deliver excellent results.
He firmly believes that automation is the key to successfully embracing cloud computing, as automation unlocks the full cloud advantages. It also ensures processes are repeatable, which is a foundation of cost-optimised architectures.
His ability to take complex topics and explain them simply is one of the reasons he is part of the Microsoft MVP program. He also holds numerous certifications across all the major hyperscalers and cloud-native technologies.