Back to job search
UO
University of VictoriaVerified Job Source

Senior Advanced Research Computing Systems Administrator

The role involves designing, building, and maintaining the university's research servers, storage, and high-performance computing (HPC) infrastructure. It includes managing cloud environments, container orchestration, and providing 24/7 operational support for critical research systems.

  • On-site
  • Victoria, BC
  • Posted Jul 25, 2026
  • Apply by Aug 27, 2026
  • 1 position

More jobs you can apply to directly

Similar opportunities posted by employers hiring on Jobs.ca, with no external application form.

Job summary

Mandate Reporting to the Manager and Architect of Advanced Research Computing Infrastructure, the Senior Advanced Research Computing Systems Administrator works as part of a team to design, build, and ensure the operational effectiveness of the university’s research servers and storage. Members of this team maintain systems critical to many research groups on-campus and beyond, including web servers and database servers, and large, high-performance research computing systems (HPC),cloud infrastructure and container orchestration used by researchers both at UVic, from institutions across the country, and with international collaborations. These systems are required to be in operation 24 hours per day, 365 days of the year and decisions regarding these systems can impact UVic’s obligations to other parties beyond the institution. Objectives The Senior Advanced Research Computing System Administrator’s work includes the design, installation, configuration, and maintenance of hardware and software, problem determination/resolution, resource allocation, performance and security monitoring, and usage reporting. Each position has specialized areas of expertise in multiple domains storage technologies such as Ceph, dCache, GPFs, Lustre and IBM Spectrum Protect (TSM); deployment technologies like, xCAT, Cobbler, Ansible, Puppet, and Terraform; and compute/virtualization technologies such as Kubernetes, OpenStack; HPC Schedulers such as SLURM, HTCondor, Moab; and Systems Monitoring. The specific technologies that are leveraged in this role will change over time and this position has the responsibility to help guide the decision on how future technologies are selected and deployed. This position requires the incumbent to have significant problem solving skills to analyze and correct software and hardware problems and to automate administration tasks. This includes unanticipated and unique problem solutions where the incumbent may be the sole expert in the area. The incumbent also must possess effective communications skills in order to provide technical assistance and advice to peers and the user community, as well as inform user areas on the impact and implications of system failures, maintenance, and cyber security incidents. This role leads project teams and provides recommendations on the university’s server and storage infrastructure. System maintenance is usually required to be performed off-hours and major issues are responded to on a 24/7 basis. This role may need to work outside of normal work hours on an emergency or pre-scheduled basis. The role may need to travel out of town/country. This position requires a Bachelor’s Degree in Computer Science or other relevant discipline plus at least five years of experience in system administration in a large enterprise or academic/research environment. An equivalent combination of education and experience may be considered. Required Knowledge, Skills, And Abilities Include Expert knowledge of RedHat Enterprise Linux and/or derivatives (eg AlmaLinux, Rocky Linux, etc) In-depth experience installing and operating of at least one of OpenStack, Kubernetes, or Ceph In-depth experience with scripting and revision control (e.g. Bash, PERL and Python, Git or Subversion) Working knowledge of provisioning and configuration management tools (e.g. Ansible, Terraform, xCAT, Cobbler) Experience supporting cloud computing and/or containerized environments Excellent communication skills, both written and verbal Ability to build and maintain productive working relationships with all stakeholders Ability to work collaboratively in a team environment Proven track record achieving project goals on time and produce deliverables of a high quality High degree of attention to detail is required, as is the ability to understand complex technical concepts and the need to maintain broad and in-depth technical knowledge of all aspects of servers and server operating systems. High level of problem solving abilities; must be able to effectively identify and resolve unusual and highly complex technical problems. Ability to effectively manage multiple tasks and priorities and work under pressure to meet time sensitive and mission critical deadlines in a complex environment. Ability to take initiative and work with limited direction. Ability to mentor and coach technical staff and teams, and act as a resource. Ability to successfully contribute to complex projects: developing project work plans; monitoring and directing the activities of a project team. Excellent written and oral communications skills. Ability to collaborate, build and maintain positive relationships with diverse individuals and work effectively in a team environment. Commitment to valuing diversity and contributing to an inclusive and respectful working and learning environment. Assets Or Preferences Working knowledge of Load Balancers and HA environments Experience supporting HPC environments Experience supporting compute and/or storage systems in a research or academic setting Experience participating with and contributing to open-source software projects Working knowledge of GPU acceleration of computational workloads, preferably in a virtualized environment Working knowledge of KVM/QEMU virtualization, ContainerD or Docker container runtimes, and Calico, Linuxbridge, or OpenVSwitch virtual networking

What you’ll do

The role involves designing, building, and maintaining the university's research servers, storage, and high-performance computing (HPC) infrastructure. It includes managing cloud environments, container orchestration, and providing 24/7 operational support for critical research systems.

Requirements

Requires a Bachelor's degree in Computer Science or a related field with at least five years of experience in large-scale enterprise or academic system administration. Candidates must have expert knowledge of Linux and in-depth experience with virtualization, scripting, and configuration management tools.

Listed skills

  • KubernetesPreferred
  • TerraformPreferred
  • GitPreferred
  • PythonPreferred

Other relevant skills

Identified from the job description. Confirm important requirements above.

  • RedHat Enterprise Linux
  • OpenStack
  • Kubernetes
  • Ceph
  • Bash
  • Python
  • Git
  • Ansible
  • Terraform
  • SLURM
  • HPC
  • Container Orchestration
  • System Administration
  • Project Leadership
  • Network Virtualization
  • Storage Technologies

Job areas

  • Technology
  • Science & Research
  • Education
  • Software
  • Engineering

Additional details

Minimum education
Bachelor’s degree
Minimum experience
5+ years
Apply by
Aug 27, 2026
Posting language
English
Working hours
40 hours per week
Seniority
Mid-Senior level