Back to job search
E
EleksVerified Job Source

Infrastructure/GPU Cluster/Platform Operations Lead

Lead the design and operation of GPU infrastructure and Kubernetes-based AI platform environments. Optimize infrastructure for AI training and inference while defining operational standards and reliability practices.

  • Hybrid
  • Canada
  • Posted Jun 25, 2026
  • 1 position

Job summary

ELEKS is looking for an Infrastructure/GPU Cluster/Platform Operations Lead in Canada. Alberta-based candidates are strongly preferred (Calgary or Edmonton). Canada-based candidates will also be considered. ABOUT CLIENT Our customer is building a next-generation AI platform that enables organizations to securely develop, govern, and operationalize artificial intelligence while ensuring that sensitive data and organizational knowledge remain fully under their control. The platform combines advanced AI capabilities with enterprise-grade governance, security, and data sovereignty to support mission-critical decision-making. The solution serves government organizations and enterprise customers operating in highly regulated and security-sensitive environments, where reliability, accountability, and trust are essential. The platform supports intelligent decision-making across strategic planning, workforce intelligence, and organizational operations, helping customers leverage AI without compromising security, compliance, or control over their data. \n REQUIREMENTS 8+ years of Infrastructure Engineering or Platform Operations experience Experience managing GPU clusters for AI workloads Strong Kubernetes administration skills Experience with NVIDIA GPU technologies and CUDA ecosystem Experience with cloud infrastructure (Azure, AWS or GCP) Knowledge of storage, networking, and high-performance computing environments Experience implementing Infrastructure as Code (Terraform or similar) Strong operational leadership skills Experience supporting AI platform infrastructure Upper-Intermediate or higher level of English RESPONSIBILITIES Lead GPU infrastructure design and operations Manage Kubernetes-based AI platform environments Optimize infrastructure for AI training and inference workloads Define operational standards and reliability practices Collaborate with AI engineering teams Implement monitoring, security, and disaster recovery strategies Lead infrastructure capacity planning Support technical roadmap and infrastructure evolution \n

What you’ll do

Lead the design and operation of GPU infrastructure and Kubernetes-based AI platform environments. Optimize infrastructure for AI training and inference while defining operational standards and reliability practices.

Requirements

Requires over 8 years of experience in infrastructure engineering or platform operations with a strong focus on GPU clusters and Kubernetes. Proficiency in NVIDIA technologies, cloud platforms, and Infrastructure as Code is essential.

Other relevant skills

Identified from the job description. Confirm important requirements above.

  • GPU Cluster Management
  • Kubernetes Administration
  • NVIDIA GPU Technologies
  • CUDA Ecosystem
  • Cloud Infrastructure
  • Infrastructure as Code
  • Terraform
  • High-Performance Computing
  • Infrastructure Engineering
  • Platform Operations
  • Capacity Planning
  • Disaster Recovery

Job areas

  • Technology
  • Engineering
  • Software
  • Management & Leadership
  • Data & Analytics

Additional details

Minimum experience
10+ years
Posting language
English
Working hours
40 hours per week