Sessions1 - 1 of 6
Everything on the program. Search by session title or speaker, or narrow it down by track.
- InfrastructureTalk
Serving four hundred models on one GPU fleet
Multi-tenant inference from the operator's seat: scheduling, memory packing, cold-start mitigation, noisy-neighbour isolation, and the observability you need before you can safely oversubscribe anything.
Monday, October 12: 11:15 AM - 12:00 PMMain StageSpeaker
- LFLiam Ferguson
Staff Site Reliability Engineer · Orbit Cloud
Format: TalkTrack: Infrastructure - LF