Back to job search
Axiom Global Technologies logo

Kubernetes (K8s) / HPC Engineer

  • On-site
  • Posted Oct 8, 2026
  • 1 position

Opens LinkedIn

Sign in to save this job
Employment type
Contract
Experience level
Senior · 5+ years
Posting language
English
Working hours
40 hours per week
Seniority
Mid-Senior level
Application method
Direct apply is available

Job summary

Design, deploy, optimize, and manage Kubernetes clusters, HPC environments, and cloud-native compute platforms for AI/ML, engineering simulations, and data-intensive workloads. Configure GPU infrastructure, automate and monitor systems, maintain CI/CD pipelines, and troubleshoot performance and reliability issues.

Job details

Job Summary We are seeking an experienced Kubernetes (K8s) / High-Performance Computing (HPC) Engineer to design, deploy, optimize, and manage scalable compute platforms supporting AI/ML, engineering simulations, and data-intensive workloads. The ideal candidate will have strong expertise in Kubernetes, containerization, Linux administration, HPC environments, and cloud-native infrastructure, with hands-on experience supporting distributed computing and GPU-accelerated workloads. Key Responsibilities Design, deploy, and manage Kubernetes clusters and containerized applications using Docker, Helm, and related orchestration technologies. Support HPC clusters and distributed computing environments using workload schedulers such as Slurm and Volcano. Configure and optimize GPU infrastructure, particularly NVIDIA-based environments, for AI/ML workloads and high-performance computing. Administer Linux-based systems and develop automation scripts using Python and Bash. Implement and maintain CI/CD pipelines, infrastructure automation, monitoring, and performance optimization solutions. Deploy and manage cloud-native infrastructure on platforms such as Azure AKS, AWS EKS, and Google Kubernetes Engine (GKE). Troubleshoot infrastructure and workload performance issues to ensure platform reliability, scalability, and efficiency. Required Qualifications 5+ years of experience in Kubernetes and cloud infrastructure engineering. At least 2+ years of hands-on experience with HPC or large-scale computing platforms. Strong knowledge of Kubernetes, Docker, Helm, and container orchestration. Experience with HPC clusters, workload schedulers such as Slurm or Volcano, and distributed computing. Hands-on experience with NVIDIA GPU infrastructure, AI/ML workloads, and performance tuning. Strong Linux administration and Python/Bash scripting skills. Experience with CI/CD, infrastructure automation, monitoring, and cloud platforms such as Azure AKS, AWS EKS, or GCP GKE. Preferred Qualifications Knowledge of high-speed networking technologies, including RDMA and InfiniBand. Experience with parallel storage systems and large-scale compute environments. Background supporting engineering simulations, semiconductor workloads, AI model training, or scientific computing platforms. Additional Requirement Candidates must be available to start immediately and successfully complete the mandatory background check. Interested candidates: Please share your updated resume, current location in Canada, availability to start, and relevant Kubernetes/HPC experience.

What you’ll do

Design, deploy, optimize, and manage Kubernetes clusters, HPC environments, and cloud-native compute platforms for AI/ML, engineering simulations, and data-intensive workloads. Configure GPU infrastructure, automate and monitor systems, maintain CI/CD pipelines, and troubleshoot performance and reliability issues.

Requirements

Requires 5+ years of Kubernetes and cloud infrastructure engineering experience, including at least 2 years supporting HPC or large-scale computing platforms. Candidates should have hands-on experience with container orchestration, schedulers, NVIDIA GPUs, Linux, Python or Bash, CI/CD, automation, monitoring, and a major cloud Kubernetes platform; high-speed networking and parallel storage experience are preferred.

Listed skills

  • Kubernetes · Preferred
  • Docker · Preferred
  • Python · Preferred
  • CI/CD · Preferred

Other relevant skills

Identified from the job description. Confirm important requirements above.

  • Kubernetes
  • Docker
  • Helm
  • High-Performance Computing
  • Slurm
  • Volcano
  • NVIDIA GPU Infrastructure
  • Linux Administration
  • Python
  • Bash
  • CI/CD
  • Infrastructure Automation
  • Monitoring
  • Azure AKS
  • AWS EKS
  • Google Kubernetes Engine

Job areas

  • Technology
  • Software
  • Engineering
  • Data & Analytics
  • Science & Research

More jobs from Axiom Global Technologies

See all jobs from Axiom Global Technologies