Back to job search
High Tech Genesis logo
High Tech GenesisVerified Job Source

Systems Reliability Engineer

  • Toronto, ON
  • On-site
  • Posted Oct 4, 2026
  • 1 position

$50–$60 / hour

Opens an external site

Sign in to save this job
Employment type
Contract
Experience level
Senior · 5+ years
Apply by
Nov 1, 2026
Posting language
English
Working hours
40 hours per week
Seniority
Associate
Application method
Direct apply is available

Job summary

Own the reliability, resiliency, availability, and performance of large-scale enterprise platforms by building monitoring and alerting frameworks, defining SLOs, SLIs, and error budgets, and proactively addressing operational risks. Lead incident response and root cause analysis, coordinate remediation across teams, manage incident tracking, and develop automation to improve operational efficiency and reduce detection and resolution times.

Job details

Job Overview We are seeking a Systems Reliability Engineer (SRE) to join a Technical Operations team supporting a large-scale enterprise platform. In this role, you will be responsible for the reliability, resiliency, availability, and performance of critical systems. You will design and maintain monitoring and alerting solutions, lead incident response activities, drive root cause analysis, and develop automation to improve operational efficiency and platform stability. The ideal candidate has strong experience operating enterprise-scale environments and a solid understanding of reliability engineering and observability practices. Responsibilities Own the reliability, resiliency, and availability of large-scale enterprise platforms, proactively identifying and mitigating risks to service continuity. Design, implement, and maintain comprehensive monitoring, observability, and alerting frameworks using tools such as Splunk, Dynatrace, Grafana, and Datadog. Define and track Service Level Objectives (SLOs), Service Level Indicators (SLIs), and error budgets to measure and improve platform health. Lead and participate in incident response activities, serving as a technical driver during remediation efforts. Coordinate with engineering, product, operations, and risk teams during incidents and service-impacting events. Lead and improve the Root Cause Analysis (RCA) process by investigating incidents, documenting events and remediation actions, and identifying underlying causes to prevent recurrence. Ensure timely creation, management, and tracking of incident tickets using platforms such as ServiceNow. Monitor incident aging, trends, and reporting to identify opportunities for operational improvements. Build automation and operational tooling to reduce manual effort and improve Mean Time to Detection (MTTD) and Mean Time to Resolution (MTTR). Collaborate with engineering and product teams to incorporate reliability and resiliency best practices throughout the platform lifecycle. Proactively identify opportunities to improve system performance, availability, monitoring, and operational efficiency. Required Qualifications Hands-on experience with monitoring, observability, and alerting tools, including Splunk, Dynatrace, Grafana, and Datadog. Proven experience operating and supporting large-scale enterprise platform environments. Demonstrated experience with incident response and leading or contributing to Root Cause Analysis (RCA)processes. Strong understanding of Site Reliability Engineering (SRE) principles, including availability, resiliency, monitoring, alerting, and system performance. Experience with incident management and ticketing workflows, such as ServiceNow. Strong troubleshooting and problem-solving skills. Excellent communication skills with the ability to coordinate technical remediation efforts across multiple teams. Ability to work effectively in a fast-paced environment and respond to high-priority incidents. Preferred Qualifications Experience working within financial services, payments, or embedded finance environments. Proficiency with scripting or programming languages such as Python, Go, or Bash. Familiarity with cloud platforms, containerization, and CI/CD pipelines. Experience defining and managing SLOs, SLIs, and error budgets. Experience developing automation and tooling to improve operational reliability and reduce repetitive manual tasks. Knowledge of modern cloud-native architectures and distributed systems.

What you’ll do

Own the reliability, resiliency, availability, and performance of large-scale enterprise platforms by building monitoring and alerting frameworks, defining SLOs, SLIs, and error budgets, and proactively addressing operational risks. Lead incident response and root cause analysis, coordinate remediation across teams, manage incident tracking, and develop automation to improve operational efficiency and reduce detection and resolution times.

Requirements

Requires hands-on experience with observability tools such as Splunk, Dynatrace, Grafana, and Datadog, along with experience supporting large-scale enterprise platforms and participating in incident response and root cause analysis. Candidates should understand SRE principles and incident management, have strong troubleshooting and communication skills, and be able to coordinate effectively during high-priority incidents.

Listed skills

  • Splunk · Preferred
  • Troubleshooting · Preferred
  • Root Cause Analysis · Preferred

Other relevant skills

Identified from the job description. Confirm important requirements above.

  • Site Reliability Engineering
  • Monitoring
  • Observability
  • Alerting
  • Incident Response
  • Root Cause Analysis
  • Service Level Objectives
  • Service Level Indicators
  • Error Budgets
  • Enterprise Platform Operations
  • Troubleshooting
  • Automation
  • Splunk
  • Dynatrace
  • Grafana
  • Datadog

Job areas

  • Technology
  • Software
  • Engineering

More jobs from High Tech Genesis

See all jobs from High Tech Genesis