- MLOps North
- Toronto Summit
- Presented by TMLS
Evals and Benchmarks
About the track
In a world where models are released daily, how do we know when the right time is to swap models? External benchmarks give us hints, but internal, use-case-specific ones help teams make decisions on accuracy, cost and latency. This track is for ML/AI engineers looking to grow their intermediate-to-advanced eval skills.
Talks cover industry-leading benchmarks, how to build internal benchmarks, when evals fail to lead to business outcomes, and auditing benchmarks. Expect practical eval advice that will help teams move more quickly, build better products and save on costs. You’ll walk out with a blueprint for building an effective eval.
Track host
VP AI Engineering, Vector Institute
See Track Lead Profile
Kathryn Hume leads AI Engineering at the Vector Institute, where she builds tools and platforms that further industry adoption of frontier AI. Prior to joining Vector, she held multiple technology leadership roles in financial services and AI, including Chief Technology Officer for Tangerine Bank, Vice President of Technology for RBC’s mobile, online banking, and direct investing businesses, and Interim Head of Borealis AI, RBC’s AI research institute. Before joining RBC, she held leadership roles at AI startups integrate.ai and Fast Forward Labs (acquired by Cloudera). Kathryn is a recognized author and public speaker, with work appearing at TED, the Harvard Business Review, and the Globe & Mail. She has served as a guest lecturer and adjunct professor — teaching courses on applied AI and responsible AI — at the Harvard Business School, Stanford, the University of Toronto, and the University of Calgary. She is a mother of two young boys, a board member at CanadaHelps active in the charity sector, and an advocate for women in technology and educational reform in the age of generative AI.
Toronto Summit · Presented by TMLS · Nov 5–6, 2026 · RBC WaterPark Place, TorontoThe people building agents, and the people building with them – one room, on the Toronto waterfront. Two days of practitioner-curated talks on what actually ships: the stack underneath agentic systems, and the products, workflows, and agents teams are running in production right now.