Back to job search
Motion Recruitment logo

Staff Site Reliability Engineer/ Azure

  • Toronto, ON
  • On-site
  • Posted Oct 9, 2026
  • 1 position

Opens LinkedIn

Sign in to save this job
Employment type
Full-time
Experience level
Lead · 10+ years
Posting language
English
Working hours
40 hours per week
Seniority
Mid-Senior level
Application method
Direct apply is available

Job summary

Own platform reliability and infrastructure strategy across Azure, .NET, legacy applications, and modern cloud systems. Improve observability, automation, incident management, and system performance while mentoring engineers and collaborating across DevOps and product engineering teams.

Job details

Join a growing financial technology organization where your engineering expertise will make a meaningful difference in how schools across North America manage their financial operations. As a Staff Site Reliability Engineer, you'll have the opportunity to take real ownership of platform reliability, shape the future of infrastructure strategy, and help build resilient systems that thousands of educational institutions depend on every day. Working across Microsoft Azure, .NET, legacy applications, and modern cloud technologies, you'll tackle interesting technical challenges while driving improvements in observability, automation, incident management, and overall system performance. You'll also have the freedom to explore emerging AI-driven solutions, collaborate with talented DevOps and engineering teams, and mentor others as the organization continues to grow and evolve. If you're someone who loves solving complex problems, enjoys staying close to the technology, and wants the opportunity to leave a lasting mark on both the platform and the engineering culture, this is an exciting opportunity to do exactly that. Required Skills & Experience 10+ years of experience in Site Reliability Engineering, Platform Engineering, or infrastructure operations, with proven expertise establishing reliability strategies, defining SLIs/SLOs and error budgets, leading production incident response, and implementing effective root cause analysis and postmortem practices. Strong hands-on experience with Microsoft Azure, .NET environments, legacy .NET Framework applications, and IIS, including the ability to improve availability, resilience, and operational performance across both modern and established production systems. Advanced expertise building observability and monitoring capabilities across metrics, logging, distributed tracing, and alerting, combined with strong scripting and automation skills to eliminate operational toil, improve incident detection, and develop maintainable production-grade tooling. Desired Skills & Experience Advanced experience in capacity planning, performance optimization, load testing, and proactive infrastructure scaling, with the ability to identify system bottlenecks and prevent reliability issues before they affect production services. Exposure to AI-driven operations and intelligent observability, including applying machine learning or AI-assisted tooling to anomaly detection, incident triage, operational analytics, and automated remediation workflows. Demonstrated technical leadership in coaching senior engineers, improving on-call practices, establishing reliability engineering standards, and influencing cross-functional architecture and operational decisions across DevOps and product engineering teams. Daily Responsibilities Hands-On Engineering: 70% Team Collaboration & Cross-Functional Work: 30% You Will Receive The Following Benefits Medical, Dental, and Vision Insurance Vacation Time Current Vacancy: Yes Use of AI in Hiring: No Applicants must be currently authorized to work in Canada on a full-time basis now and in the future. Accommodation will be provided in all parts of the hiring process as required under Motion Recruitment’s Employment Accommodation policy. Applicants need to make their needs known in advance. Posted By: Adrian Cronk

What you’ll do

Own platform reliability and infrastructure strategy across Azure, .NET, legacy applications, and modern cloud systems. Improve observability, automation, incident management, and system performance while mentoring engineers and collaborating across DevOps and product engineering teams.

Requirements

Requires 10+ years in site reliability engineering, platform engineering, or infrastructure operations, including experience defining reliability strategies, SLIs/SLOs, and error budgets and leading incident response and postmortems. Candidates need hands-on Azure, .NET, legacy .NET Framework, and IIS experience, plus advanced observability, monitoring, scripting, and automation skills.

Benefits

  • Medical Insurance
  • Dental Insurance
  • Vision Insurance
  • Vacation Time

Listed skills

  • Microsoft Azure · Preferred
  • .NET · Preferred
  • Root Cause Analysis · Preferred

Other relevant skills

Identified from the job description. Confirm important requirements above.

  • Site Reliability Engineering
  • Microsoft Azure
  • .NET
  • .NET Framework
  • IIS
  • Reliability Engineering
  • SLIs and SLOs
  • Error Budgets
  • Incident Response
  • Root Cause Analysis
  • Observability
  • Monitoring
  • Distributed Tracing
  • Scripting and Automation
  • Capacity Planning
  • Performance Optimization

Job areas

  • Technology
  • Software
  • Engineering
  • Management & Leadership

More jobs from Motion Recruitment

See all jobs from Motion Recruitment