AI Engineer Summit 2026

Oct 12–13, 2026Moscone West, San Francisco

Back to all sessions
InfrastructureTalk

Serving four hundred models on one GPU fleet

Monday, October 12, 2026: 11:15 AM - 12:00 PMMain Stage · seats 300

About this session

Multi-tenant inference from the operator's seat: scheduling, memory packing, cold-start mitigation, noisy-neighbour isolation, and the observability you need before you can safely oversubscribe anything.

Speaker

LF

Liam Ferguson

Staff Site Reliability Engineer

Orbit Cloud

Format: TalkTrack: InfrastructureLevel: AdvancedLanguage: Englishgpuoperations