Site Reliability Engineering Manager
- Toronto, ON
- Hybrid
- Posted Sep 18, 2026
- 1 position
Opens an external site
- Employment type
- Contract
- Experience level
- Senior · 6+ years
- Apply by
- Oct 15, 2026
- Posting language
- English
- Working hours
- 40 hours per week
- Seniority
- Mid-Senior level
- Application method
- Direct apply is available
Job summary
Responsible for day-to-day reliability and the implementation of Service Level Availability (SLAs) and reliability metrics. Manages high-availability production environments, incident response, and service restoration using AI-powered engineering tools.
Job details
Experienced Application Manager (Site Reliability Engineer) who is responsible for day-to-day reliability Implement and operationalize Service Level Availability (SLAs) and related reliability metrics 6+ years of experience in SRE, Production Engineering, Platform Engineering, Application Support, or Technology Service Delivery. Strong technical knowledge of cloud platforms, enterprise systems, and application, data, and platform architectures. Capital Markets Total Fund Market Investment Proven experience managing high-availability production environments, incident response, troubleshooting, and root cause analysis. AI-powered engineering tools, including IDE assistants and MCP-enabled integrations, to support troubleshooting and service restoration. Experience with observability, automation, Python, and PowerShell. Experience working with APIs, messaging/queueing platforms, distributed systems, CI/CD, Infrastructure as Code, DevOps, and DataOps. Working knowledge of Agile, Waterfall, DevOps, ITIL, and COBIT. Experience with Jira, Confluence, and Git. Excellent communication and collaboration skills, with the ability to partner effectively with architects, business analysts, DBAs, QA, and cross-functional technology teams. Contract/Hybrid Toronto
What you’ll do
Responsible for day-to-day reliability and the implementation of Service Level Availability (SLAs) and reliability metrics. Manages high-availability production environments, incident response, and service restoration using AI-powered engineering tools.
Requirements
Requires 6+ years of experience in SRE or related platform engineering roles with strong knowledge of cloud platforms and distributed systems. Proficiency in Python, PowerShell, and various DevOps/DataOps methodologies is essential.
Listed skills
- CI/CD · Preferred
- Root Cause Analysis · Preferred
- Agile · Preferred
- Python · Preferred
Other relevant skills
Identified from the job description. Confirm important requirements above.
- SRE
- Production Engineering
- Cloud Platforms
- Incident Response
- Root Cause Analysis
- Observability
- Automation
- Python
- PowerShell
- APIs
- CI/CD
- Infrastructure as Code
- DevOps
- DataOps
- Agile
- ITIL
Job areas
- Technology
- Software
- Engineering
- Management & Leadership
- Finance & Accounting
More jobs you can apply to directly
Similar opportunities posted by employers hiring on Jobs.ca, with no external application form.
Bédard Ressources Humaines
ITAD Services Representative #1265
SponsoredDirect employerEasy Apply- On-site
- Mississauga, ON
- Posted Sep 16, 2026
Sustainable Projects Group
Sustainability consultant/ Team Lead, Energy Consulting NOC 41400
SponsoredDirect employerEasy Apply- Hybrid
- vancouver v5l 4s1, BC
- Posted Sep 17, 2026
Desjardins
Senior Litigation Advisor -
SponsoredDirect employerEasy Apply- Hybrid
- Calgary, AB
- Posted Sep 9, 2026
