Building with Agents

Evals and Benchmarks

About the track

In a world where models are released daily, how do we know when the right time is to swap models? External benchmarks give us hints, but internal, use-case-specific ones help teams make decisions on accuracy, cost and latency. This track is for ML/AI engineers looking to grow their intermediate-to-advanced eval skills.

Talks cover industry-leading benchmarks, how to build internal benchmarks, when evals fail to lead to business outcomes, and auditing benchmarks. Expect practical eval advice that will help teams move more quickly, build better products and save on costs. You’ll walk out with a blueprint for building an effective eval.

Track host

Kathryn Hume

VP AI Engineering, Vector Institute

Kathryn Hume leads AI Engineering at the Vector Institute, where she builds tools and platforms that further industry adoption of frontier AI. Prior to joining Vector, she held multiple technology leadership roles in financial services and AI, including Chief Technology Officer for Tangerine Bank, Vice President of Technology for RBC’s mobile, online banking, and direct investing businesses, and Interim Head of Borealis AI, RBC’s AI research institute. Before joining RBC, she held leadership roles at AI startups integrate.ai and Fast Forward Labs (acquired by Cloudera). Kathryn is a recognized author and public speaker, with work appearing at TED, the Harvard Business Review, and the Globe & Mail. She has served as a guest lecturer and adjunct professor — teaching courses on applied AI and responsible AI — at the Harvard Business School, Stanford, the University of Toronto, and the University of Calgary. She is a mother of two young boys, a board member at CanadaHelps active in the charity sector, and an advocate for women in technology and educational reform in the age of generative AI.

Summit (2 days).

Toronto Summit · Presented by TMLS · Nov 5–6, 2026 · RBC WaterPark Place, TorontoThe people building agents, and the people building with them – one room, on the Toronto waterfront. Two days of practitioner-curated talks on what actually ships: the stack underneath agentic systems, and the products, workflows, and agents teams are running in production right now.

Share

Who Attends

Attendees
0 +
Data Practitioners
0 %
Researchers/Academics
0 %
Business Leaders
0 %

2023 Event Demographics

Technical practitioners working directly with ML/AI systems
0 %
Currently Working in Industry*
0 %
Attendees Looking for Solutions
0 %
Currently Hiring
0 %
Attendees Actively Job-Searching
0 %

2023 Technical Background

Expert/Researcher
14%
Advanced
37%
Intermediate
28%
Beginner
7%

2023 Attendees & Thought Leadership

Attendees
0 +
Speakers
0 +
Company Sponsors
0 +

Business Leaders: C-Level Executives, Project Managers, and Product Owners will get to explore best practices, methodologies, principles, and practices for achieving ROI.

Engineers, Researchers, Data Practitioners: Will get a better understanding of the challenges, solutions, and ideas being offered via breakouts & workshops on Natural Language Processing, Neural Nets, Reinforcement Learning, Generative Adversarial Networks (GANs), Evolution Strategies, AutoML, and more.

Job Seekers: Will have the opportunity to network virtually and meet over 30+ Top Al Companies.

Ignite what is an Ignite Talk?

Ignite is an innovative and fast-paced style used to deliver a concise presentation.

During an Ignite Talk, presenters discuss their research using 20 image-centric slides which automatically advance every 15 seconds.

The result is a fun and engaging five-minute presentation.

You can see all our speakers and full agenda here

Get our official conference app
For Blackberry or Windows Phone, Click here
For feature details, visit Whova