Back to job search
S
SystematixVerified Job Source

Senior Site Reliability Engineer – AI & Data Platform

Establish SRE practices, standards, and observability capabilities for an enterprise AI and Data platform on Azure. Design monitoring systems, automate operational processes, and partner with DevOps and MLOps teams to ensure platform resilience and scalability.

  • Remote
  • Canada
  • Posted Aug 18, 2026
  • 1 position

More jobs you can apply to directly

Similar opportunities posted by employers hiring on Jobs.ca, with no external application form.

Job summary

We are Systematix and we are currently looking for a Senior Site Reliability Engineer – AI & Data Platform to establish the reliability, observability, and operational engineering capabilities supporting an emerging enterprise AI and Data ecosystem for one of our key clients. ABOUT THE PROJECT Our client is a global leader in science and technology, supporting a diverse portfolio of businesses across healthcare, life sciences, diagnostics, manufacturing, and industrial innovation. As investment in AI, machine learning, and advanced data capabilities continues to accelerate, the organization is building the engineering foundation required to operate these environments reliably and at enterprise scale. Current monitoring, alerting, observability, and automated operational capabilities are still evolving. The successful candidate will have an opportunity to establish modern SRE practices rather than simply inherit an existing mature environment. Working closely with Platform Engineering, DevOps, and MLOps teams, this individual will help build the standards, tooling, automation, and operating practices required to deliver highly available, resilient, and observable AI and Data platforms. ABOUT THE RESPONSIBILITIES Design and implement monitoring, observability, and alerting capabilities across Microsoft Azure cloud infrastructure and AI/ML environments. Establish Site Reliability Engineering practices, standards, operating models, and engineering principles. Define and implement Service Level Indicators (SLIs), Service Level Objectives (SLOs), and appropriate platform reliability metrics. Build dashboards, telemetry, and actionable alerting that provide meaningful visibility into platform health and performance. Implement observability solutions using technologies such as Prometheus, Grafana, OpenTelemetry, or comparable platforms. Automate operational processes, remediation, recovery, and repetitive support activities wherever possible. Improve platform availability, scalability, resilience, performance, and recoverability. Develop proactive infrastructure health, capacity, and performance monitoring capabilities. Support the reliability of production AI/ML applications, services, and GPU-intensive workloads. Partner with Platform Engineering teams to incorporate reliability and resilience into cloud infrastructure architecture. Partner with DevOps engineers to improve deployment reliability and automate operational processes. Partner with MLOps engineers to establish monitoring and observability for model, inference, and supporting infrastructure. Develop incident response processes, runbooks, operational tooling, and automated recovery capabilities. Support capacity planning and performance management across cloud and AI infrastructure. Lead and contribute to root-cause analysis and post-incident reviews. Identify recurring operational issues and develop engineering solutions that eliminate or significantly reduce manual intervention. Drive continuous improvement of platform reliability, observability, and operational efficiency. ABOUT THE REQUIREMENTS Extensive senior-level experience in Site Reliability Engineering, Production Engineering, Systems Engineering, or a closely related engineering discipline. Strong hands-on experience with Microsoft Azure and enterprise cloud infrastructure. Deep expertise in monitoring, observability, telemetry, and alerting. Hands-on experience with Prometheus, Grafana, OpenTelemetry, or comparable observability technologies. Demonstrated experience defining and implementing SLIs, SLOs, reliability metrics, and actionable alerting. Strong Infrastructure as Code experience. Advanced automation and scripting capabilities. Strong experience with Docker, Kubernetes, and containerized production environments. Experience with CI/CD pipelines and modern software delivery practices. Demonstrated experience operating and improving complex production systems at enterprise scale. Experience automating operational processes, remediation, and recovery activities. Strong understanding of distributed systems, cloud architecture, scalability, resilience, and performance. Exceptional troubleshooting, root-cause analysis, and systems-thinking capabilities. Strong software engineering mindset with a demonstrated preference for solving operational problems through engineering and automation rather than manual processes. Strong communication and collaboration skills with the ability to work across Platform Engineering, DevOps, MLOps, cloud, security, data, and application engineering teams. PREFERRED QUALIFICATIONS Experience supporting AI and machine learning infrastructure in production environments. Experience supporting GPU, accelerated computing, or high-performance computing environments. Strong Python development or scripting experience. Experience with Azure Machine Learning and related Azure AI/data services. Experience with MLOps platforms and machine learning lifecycle infrastructure. Experience implementing observability for machine learning models, inference services, and supporting infrastructure. Experience establishing an SRE capability, standards, or operating model within an immature or evolving engineering environment. Experience working within large, global, or federated enterprise environments. ABOUT THE ROLE This is a remote contract opportunity supporting a strategic enterprise AI, Data, and Cloud engineering initiative. The successful candidate must be based within, or able to consistently work, Eastern Time Zone business hours.This is a senior, highly hands-on engineering position. The successful candidate will help establish the organization's SRE capability while personally designing and implementing the observability, automation, reliability, and operational tooling required to support the platform. We are looking for a software and engineering-oriented SRE rather than a traditional operations or support professional. The ideal candidate approaches recurring operational problems as opportunities to engineer and automate them away. AI DISCLOSURE As part of our recruitment process, Systematix may use artificial intelligence (AI) tools to assist with resume screening, candidate matching, and recruitment administration. All hiring decisions are ultimately made by our recruitment and hiring teams. APPLY NOW If you are interested in finding out more, please contact us or submit your resume to [email protected]. Know someone who would be a great fit? We welcome referrals of qualified candidates and are always interested in connecting with talented technology professionals. ABOUT SYSTEMATIX Systematix is a Canadian-owned Global Consulting and Resourcing firm with nearly 50 years of experience delivering technology solutions to clients across North America and the United Kingdom. We provide the highest-caliber consulting solutions to a diverse client base across all levels of government and private industry. Systematix is committed to creating a diverse, inclusive environment and is proud to be an equal opportunity employer. At Systematix, we value diverse perspectives, experiences, and backgrounds. Systematix. Solutions Focused. People Driven.

What you’ll do

Establish SRE practices, standards, and observability capabilities for an enterprise AI and Data platform on Azure. Design monitoring systems, automate operational processes, and partner with DevOps and MLOps teams to ensure platform resilience and scalability.

Requirements

Requires extensive senior-level experience in SRE or Systems Engineering with deep expertise in Azure and observability tools like Prometheus and Grafana. Candidates must have strong skills in Kubernetes, Infrastructure as Code, and a software engineering mindset for automating operational tasks.

Listed skills

  • Microsoft AzurePreferred
  • KubernetesPreferred
  • CI/CDPreferred
  • DockerPreferred
  • PythonPreferred

Other relevant skills

Identified from the job description. Confirm important requirements above.

  • Site Reliability Engineering
  • Microsoft Azure
  • Prometheus
  • Grafana
  • OpenTelemetry
  • Infrastructure as Code
  • Docker
  • Kubernetes
  • CI/CD
  • Python
  • MLOps
  • SLIs/SLOs
  • Distributed Systems
  • Root Cause Analysis
  • Observability
  • Automation

Job areas

  • Technology
  • Software
  • Data & Analytics
  • Engineering
  • Consulting

Additional details

Minimum experience
10+ years
Posting language
English
Working hours
40 hours per week
Seniority
Mid-Senior level
Application method
Direct apply is available