AI Engineer Summit 2026

Oct 12–13, 2026Moscone West, San Francisco

Back to all sessions
AI EngineeringTalk

Evaluation harnesses for production LLMs

Monday, October 12, 2026: 10:00 AM - 10:45 AMWorkshop Room · seats 80

About this session

A working tour of a real evaluation harness: dataset curation, judge prompts you can defend, statistical significance on small samples, and wiring the whole thing into CI so a regression blocks a deploy rather than surprising a customer.

Speakers

TB

Tom Beaumont

Developer Advocate

Ridgeline

HK

Hana Kobayashi

Research Engineer

Institute for Applied ML

Hana works on calibration and abstention — teaching models to recognise the edge of their own competence. She publishes regularly and reviews for three conferences.

Format: TalkTrack: AI EngineeringLevel: AdvancedLanguage: Englishevaluationci