Evals and Benchmarks
About the track
In a world where models are released daily, how do we know when the right time is to swap models? External benchmarks give us hints, but internal, use-case-specific ones help teams make decisions on accuracy, cost and latency. This track is for ML/AI engineers looking to grow their intermediate-to-advanced eval skills.
Talks cover industry-leading benchmarks, how to build internal benchmarks, when evals fail to lead to business outcomes, and auditing benchmarks. Expect practical eval advice that will help teams move more quickly, build better products and save on costs. You’ll walk out with a blueprint for building an effective eval.
Track host
Suhas Pai
Co-founder & CTO, Hudson Labs
See Track Lead Profile
Lorem Ipsum
Summit (2 days).
MLOps World | GenAI Summit 2026 is a two days of case studies, workshops, and expo on taking AI/ML and agentic systems into production – at the Etter-Harbin Alumni photo – full-bleed hero or browse files.