Back to job search
J
JobgetherVerified Job Source

Site Reliability Engineer

Design, implement, and optimize highly available infrastructure while ensuring system reliability and scalability. Proactively monitor production environments, troubleshoot complex technical issues, and manage infrastructure automation using modern tools.

  • Remote
  • Canada
  • Posted Aug 5, 2026
  • 1 position

Job summary

This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Site Reliability Engineer based in Canada. Join a technology-driven team responsible for maintaining the reliability, scalability, and security of mission-critical infrastructure supporting high-availability platforms. In this role, you will combine infrastructure engineering, automation, and operational excellence to ensure seamless system performance across cloud and on-premises environments. Working alongside cross-functional engineering teams, you will proactively optimize production systems, strengthen monitoring capabilities, and respond to complex operational challenges. This is an excellent opportunity for an experienced infrastructure professional who enjoys solving technical problems, improving reliability through automation, and contributing to resilient, enterprise-grade services in a fast-paced environment. \n Accountabilities Design, implement, maintain, and optimize highly available infrastructure supporting mission-critical applications and services. Monitor production environments, analyze system performance, and proactively identify opportunities to improve stability, scalability, and operational efficiency. Respond to technical escalations, troubleshoot infrastructure, networking, hardware, and software issues, and lead resolution of critical incidents. Develop and maintain monitoring, alerting, backup, recovery, and disaster recovery procedures to maximize uptime and minimize business impact. Manage cloud platforms, virtualization technologies, network infrastructure, and remote monitoring systems to ensure secure and reliable operations. Build and maintain infrastructure automation using configuration management, scripting, and Infrastructure-as-Code tools. Participate in post-incident reviews, document operational improvements, and contribute to continuous reliability and security enhancements. Collaborate with engineering and operations teams to strengthen CI/CD pipelines, system resilience, and infrastructure best practices. Requirements Bachelor's degree in Computer Science, Information Technology, or a related field; relevant professional certifications are advantageous. At least 3 years of experience as a Site Reliability Engineer, DevOps Engineer, Systems Administrator, or in a similar infrastructure role. Strong experience with cloud platforms such as AWS or Oracle Cloud. Proficiency with Linux and Windows server administration, virtualization technologies, and enterprise infrastructure management. Experience with Docker, Kubernetes, Terraform, Git, GitLab CI/CD, ELK Stack, Prometheus, and Grafana. Knowledge of MySQL, PostgreSQL, networking concepts (LAN/WAN, HTTP, TCP/IP), system security, and backup/recovery strategies. Experience with automation and scripting tools such as Ansible, Bash, Rundeck, or Puppet. Familiarity with Nginx, PHP-FPM, SSL, DNS, and Cloudflare is considered an advantage. Strong analytical thinking, troubleshooting skills, attention to detail, and the ability to work independently and collaboratively. Availability to respond to critical production incidents outside standard business hours when required. Benefits Competitive salary package. Private health insurance. Annual wellness allowance. Birthday leave. Company-sponsored team-building events and social activities. Relocation support, where applicable. Opportunity to work with modern cloud technologies, automation tools, and large-scale infrastructure. Professional development opportunities within a collaborative engineering environment. \n How Jobgether works: We use an AI-powered matching process to ensure your application is reviewed quickly, objectively, and fairly against the role's core requirements. Our system identifies the top-fitting candidates, and this shortlist is then shared directly with the hiring company. The final decision and next steps (interviews, assessments) are managed by their internal team. We appreciate your interest and wish you the best! Why Apply Through Jobgether? Data Privacy Notice: By submitting your application, you acknowledge that Jobgether will process your personal data to evaluate your candidacy and share relevant information with the hiring employer. This processing is based on legitimate interest and pre-contractual measures under applicable data protection laws (including GDPR). You may exercise your rights (access, rectification, erasure, objection) at any time. #LI-CL1

What you’ll do

Design, implement, and optimize highly available infrastructure while ensuring system reliability and scalability. Proactively monitor production environments, troubleshoot complex technical issues, and manage infrastructure automation using modern tools.

Requirements

Requires a bachelor's degree in Computer Science or a related field and at least 3 years of experience in infrastructure or DevOps roles. Candidates must possess strong proficiency in cloud platforms, Linux/Windows administration, and automation scripting.

Benefits

• Competitive salary package • Private health insurance • Annual wellness allowance • Birthday leave • Company-sponsored team-building events • Relocation support • Professional development opportunities

Listed skills

  • KubernetesPreferred
  • DockerPreferred
  • Amazon Web ServicesPreferred
  • LinuxPreferred
  • TerraformPreferred
  • GitPreferred

Other relevant skills

Identified from the job description. Confirm important requirements above.

  • Site Reliability Engineering
  • Cloud Platforms
  • AWS
  • Oracle Cloud
  • Linux
  • Windows Server Administration
  • Docker
  • Kubernetes
  • Terraform
  • Git
  • GitLab CI/CD
  • ELK Stack
  • Prometheus
  • Grafana
  • Ansible
  • Infrastructure-as-Code
  • Microsoft Windows Server Administration
  • Backup and Recovery System
  • Operational Efficiency
  • Pipelines
  • Network Infrastructure
  • CI/CD
  • Infrastructure Automation
  • General Data Protection Regulation (GDPR)
  • Puppet (Configuration Management Tool)
  • Analytical Thinking
  • Git (Version Control System)
  • Data Privacy
  • Resilience
  • Infrastructure as Code (IaC)
  • PHP (Scripting Language)
  • Artificial Intelligence
  • Amazon Web Services
  • Automation
  • Bash (Scripting Language)
  • Business Continuity Planning
  • Configuration Management
  • Computer Science
  • Information Technology
  • Information Privacy
  • DevOps
  • Disaster Recovery
  • Scalability
  • Infrastructure Management
  • Networking Hardware
  • Local Area Networks
  • PostgreSQL
  • MySQL
  • Nginx
  • Operational Excellence

Job areas

  • Technology
  • Software
  • Engineering
  • Security & Safety
  • Site Reliability Engineer
  • Software and Applications Developers and Analysts Not Elsewhere Classified
  • Computer and Information Research Scientists

Additional details

Minimum education
Bachelor’s degree
Minimum experience
2+ years
Posting language
English
Working hours
40 hours per week