As AI agents move from demos into real-world workflows, accuracy and reliability become critical. An agent can produce impressive results in a controlled demo, but how do we know it will make the right decision consistently, handle unexpected situations, and continue to perform reliably in production?
Join the AAIF Accuracy & Reliability Working Group for a virtual summit focused on the practical challenges of building and evaluating dependable agentic systems.
We’ll bring together practitioners and researchers to explore how agents are evaluated today, what we can actually measure, where current approaches fall short, and how teams are dealing with failure modes in real-world systems.
The session will feature talks and open discussion covering agent evaluation, benchmarks, reliability, verification, failure analysis, and practical lessons from deploying agent systems.
We’ll also dig into the harder questions: What does “reliable” actually mean for an agent? How should we measure it? What happens when existing evaluations don’t reflect real-world behavior? And what should the community be working on next?
Expect practical examples, lessons learned, and an open conversation about the challenges and opportunities in making AI agents systems that people can genuinely depend on.