InfrastructureTalk
Serving four hundred models on one GPU fleet
Monday, October 12, 2026: 11:15 AM - 12:00 PMMain Stage · seats 300
About this session
Multi-tenant inference from the operator's seat: scheduling, memory packing, cold-start mitigation, noisy-neighbour isolation, and the observability you need before you can safely oversubscribe anything.
Speaker
LF
Liam Ferguson
Staff Site Reliability Engineer
Orbit Cloud
Format: TalkTrack: InfrastructureLevel: AdvancedLanguage: Englishgpuoperations