Opens an external site
- Employment type
- Contract
- Experience level
- Mid-level · 2+ years
- Posting language
- English
- Working hours
- 40 hours per week
Job summary
Manage and maintain High Performance Computing (HPC) clusters, including hardware, networking, and job schedulers. Provide hands-on technical support to scientists for application installation, debugging, and runtime optimization.
Job details
MGIS is seeking a System Administrator, Level 2, to manage High Performance Computing (HPC) clusters and support the scientists who rely on them. This role blends HPC system administration with hands-on user support — helping researchers install, run, and debug applications on HPC infrastructure so they can focus on their science instead of IT issues. HPC environments in scope include clustered CPU/GPU systems with job schedulers and attached parallel storage (e.g., Lustre, GPFS). What you'll be doing HPC Administrator duties Maintain the HPC cluster — hardware, image management, local networking, scheduler, and backups Troubleshoot environment incidents to ensure a quick return to normal operations HPC Analyst duties Meet with scientists to evaluate their HPC support requirements Develop task plans to meet researchers' needs, consulting the technical authority for approval Support application builds, installs, and runtime troubleshooting (GNU, Intel, Fortran, Nvidia) Support open-source and commercial software, including Python/Anaconda installs, Bash scripting, build/make tools, EasyBuild, Spack, and MPI implementations (MPICH, OpenMPI, IntelMPI, HPMPI) Assist with compilation and runtime of in-house developed applications General systems management Manage Linux OS patching schedules and reliability Manage user accounts (creation, deletion) and environment modules Manage configuration via Git, MS DevOps, and Ansible Playbooks Manage RPM/DEB packages and troubleshoot ThinLinc Troubleshooting & hardware Troubleshoot jobs on schedulers (PBS Pro/Torque, SLURM, SGE) Ensure reliable CUDA installs; troubleshoot GPU failures and CUDA software/driver issues Provide hardware support — memory upgrades, storage arrays, power/network cabling, ILO Documentation Document every process and task to support enterprise knowledge continuity Submit weekly progress reports to the Technical Authority Requirements What we're looking for Solid experience administering Linux-based HPC clusters (CPU/GPU nodes, schedulers, parallel storage) Hands-on experience with job schedulers such as PBS Pro/Torque, SLURM, or SGE Experience troubleshooting CUDA installations, GPU failures, and driver issues Familiarity with scientific computing toolchains — compilers (GNU, Intel), MPI implementations, EasyBuild, and Spack Experience supporting researchers or end-users with application builds and runtime issues Working knowledge of configuration management tools (Git, Ansible, MS DevOps) Comfortable working independently and producing clear technical documentation Eligible to obtain and maintain a Secret-level security clearance
What you’ll do
Manage and maintain High Performance Computing (HPC) clusters, including hardware, networking, and job schedulers. Provide hands-on technical support to scientists for application installation, debugging, and runtime optimization.
Requirements
Requires solid experience administering Linux-based HPC clusters with CPU/GPU nodes and parallel storage. Must be proficient with job schedulers, scientific toolchains, and eligible for a Secret-level security clearance.
Listed skills
- Software · Preferred
- Technical · Preferred
- Reliability · Preferred
- management · Preferred
- Process · Preferred
- Documentation · Preferred
- Troubleshooting · Preferred
- Linux · Preferred
- Technical Support · Preferred
- Git · Preferred
- Python · Preferred
- Configuration · Preferred
Other relevant skills
Identified from the job description. Confirm important requirements above.
- HPC Administration
- Linux OS
- Job Schedulers
- CUDA
- GPU Troubleshooting
- Parallel Storage
- Ansible
- Git
- Bash Scripting
- MPI
- EasyBuild
- Spack
- Python
- MS DevOps
- System Patching
- Technical Documentation
Job areas
- Technology
- Government & Public Sector
- Science & Research
- Engineering
- Software
More jobs from MGIS
Intermediate Procurement Specialist (Supply Specialist)
- On-site
- Ottawa, ON
- Posted Aug 10, 2026
Geomatics Analyst
- Hybrid
- Ottawa, ON
- Posted Aug 5, 2026
Airspace Leader/Executive
- Hybrid
- Ottawa, ON
- Posted Aug 5, 2026
Technology Architect
- Hybrid
- Ottawa, ON
- Posted Jul 7, 2026
