Cluster Site Reliability Engineer
On-site · Toronto
Bring up and validate GPU racks, maintain regional InfiniBand fabrics, and ensure cluster reliability and SLA performance. Plan capacity across regions, lead regional incident response, maintain bring-up documentation, and participate in the global pager rotation.
