
A customer-service agent can give a convincing answer and still take the wrong action. Once agents can call tools, change records or pass work between agents, evaluating the final response alone is not enough.
Using a published multi-agent customer-service reference implementation, this session will explore how to test what an agent does as well as what it says. We’ll walk through scenarios involving incorrect tool selection, missing information, failed actions and situations that require a human handoff. For each, we’ll examine what to check with deterministic tests, where model-based evaluation can help, and where human judgement remains necessary.
Attendees will leave with a practical approach to building scenario-based evaluations and a reusable release checklist for deciding whether an agent is ready for a controlled pilot. The examples will use AWS services, while the evaluation principles apply across platforms.
Olushola Oladipupo is an Enterprise Solutions Architect at Amazon Web Services, supporting customers across the UK and Ireland. He builds and publishes practical examples of generative AI and multi-agent systems, with a focus on evaluation, human oversight and enterprise adoption. He has spoken at AWS re:Invent, AWS Summits and AWS Community Day West Africa. His work connects technical architecture with the decisions teams need to make to adopt AI responsibly and effectively.