Site Reliability Engineer (SRE)
We are seeking a highly skilled Site Reliability Engineer (SRE) to join our technology team in Montreal. The ideal candidate will be responsible for improving the reliability, scalability, performance, and availability of critical enterprise applications and infrastructure. This role requires strong expertise in cloud platforms, automation, monitoring, incident management, and DevOps practices. Key Responsibilities: Design, build, and maintain highly available and scalable production systems. Implement SRE best practices including Service Level Indicators (SLIs), Service Level Objectives (SL…
- On-site
- QUEBEC
- Posted Jul 13, 2026
- Apply by Aug 12, 2026
- 1 position
Job summary
We are seeking a highly skilled Site Reliability Engineer (SRE) to join our technology team in Montreal. The ideal candidate will be responsible for improving the reliability, scalability, performance, and availability of critical enterprise applications and infrastructure. This role requires strong expertise in cloud platforms, automation, monitoring, incident management, and DevOps practices. Key Responsibilities: Design, build, and maintain highly available and scalable production systems. Implement SRE best practices including Service Level Indicators (SLIs), Service Level Objectives (SLOs), and error budgets. Automate operational tasks using scripting and Infrastructure as Code (IaC). Monitor application and infrastructure health using modern observability tools. Troubleshoot production incidents, perform root cause analysis (RCA), and implement preventive measures. Collaborate with development, infrastructure, and security teams to improve system reliability. Develop CI/CD pipelines to enable efficient and reliable software delivery. Optimize application performance, system capacity, and resource utilization. Create operational runbooks and disaster recovery procedures. Participate in on-call support and production incident management. Required Skills: 7+ years of experience in Site Reliability Engineering, DevOps, or Production Support. Strong Linux/Unix system administration experience. Hands-on experience with cloud platforms (AWS, Azure, or GCP). Strong scripting skills in Python, Bash, or Shell. Experience with Kubernetes and Docker containerization. Expertise in CI/CD tools such as Jenkins, GitHub Actions, GitLab CI, or Azure DevOps. Experience with Infrastructure as Code tools like Terraform or Ansible. Strong understanding of networking fundamentals (DNS, TCP/IP, Load Balancers, SSL). Experience with monitoring and observability tools: 1. Prometheus 2. Grafana 3. Splunk 4. ELK Stack 5. Datadog 6. AppDynamics Knowledge of logging, tracing, and metrics collection. Experience with incident management, RCA, and problem management. Understanding of high availability, disaster recovery, and capacity planning. Strong SQL and database fundamentals. Experience with Git version control.
What you’ll do
The Site Reliability Engineer will design, build, and maintain highly available and scalable production systems while implementing SRE best practices. They will also collaborate with various teams to improve system reliability and develop CI/CD pipelines for efficient software delivery.
Requirements
Candidates should have 7+ years of experience in Site Reliability Engineering or related fields, with strong Linux/Unix system administration skills. Proficiency in cloud platforms, scripting, and CI/CD tools is essential, along with a solid understanding of networking and monitoring tools.
Other relevant skills
Identified from the job description. Confirm important requirements above.
- Site Reliability Engineering
- DevOps
- Cloud Platforms
- Automation
- Monitoring
- Incident Management
- Scripting
- Infrastructure as Code
- Kubernetes
- Docker
- CI/CD
- Terraform
- Ansible
- Networking
- SQL
- Database Fundamentals
Additional details
- Minimum experience
- 5+ years
- Apply by
- Aug 12, 2026
