AI EngineeringTalk
Evaluation harnesses for production LLMs
Monday, October 12, 2026: 10:00 AM - 10:45 AMWorkshop Room · seats 80
About this session
A working tour of a real evaluation harness: dataset curation, judge prompts you can defend, statistical significance on small samples, and wiring the whole thing into CI so a regression blocks a deploy rather than surprising a customer.
Speakers
TB
Tom Beaumont
Developer Advocate
Ridgeline
HK
Hana Kobayashi
Research Engineer
Institute for Applied ML
Hana works on calibration and abstention — teaching models to recognise the edge of their own competence. She publishes regularly and reviews for three conferences.
Format: TalkTrack: AI EngineeringLevel: AdvancedLanguage: Englishevaluationci