Agentic Evaluation: Assessing Autonomous LLM Workflows
Technical Talk, Pune AI Day Q3 2026, Red Hat, Pune, Maharashtra, India
- Presented evaluation frameworks specifically designed for agentic AI systems, where standard NLP metrics fall short of capturing agent behavior.
- Detailed methodologies for assessing agent execution trajectories — evaluating the sequence of steps an agent takes, not just its final output.
- Covered tool-calling accuracy — measuring correctness, relevance, and efficiency of tool selection and parameter construction across multi-step tasks.
- Introduced end-to-end task completion metrics for autonomous LLM workflows, including partial credit scoring and failure mode taxonomy.
- Discussed hallucination risks specific to agentic contexts — compounding errors across tool calls and reasoning chains.
- Shared practical benchmarking setups and open-source frameworks used to evaluate production agentic systems at Red Hat.
