Back to job search
Bevertec logo
BevertecVerified Job Source

Site Reliability Engineering Manager

  • Toronto, ON
  • Hybrid
  • Posted Sep 18, 2026
  • 1 position

Opens an external site

Sign in to save this job
Employment type
Contract
Experience level
Senior · 6+ years
Apply by
Oct 15, 2026
Posting language
English
Working hours
40 hours per week
Seniority
Mid-Senior level
Application method
Direct apply is available

Job summary

Responsible for day-to-day reliability and the implementation of Service Level Availability (SLAs) and reliability metrics. Manages high-availability production environments, incident response, and service restoration using AI-powered engineering tools.

Job details

Experienced Application Manager (Site Reliability Engineer) who is responsible for day-to-day reliability Implement and operationalize Service Level Availability (SLAs) and related reliability metrics 6+ years of experience in SRE, Production Engineering, Platform Engineering, Application Support, or Technology Service Delivery. Strong technical knowledge of cloud platforms, enterprise systems, and application, data, and platform architectures. Capital Markets Total Fund Market Investment Proven experience managing high-availability production environments, incident response, troubleshooting, and root cause analysis. AI-powered engineering tools, including IDE assistants and MCP-enabled integrations, to support troubleshooting and service restoration. Experience with observability, automation, Python, and PowerShell. Experience working with APIs, messaging/queueing platforms, distributed systems, CI/CD, Infrastructure as Code, DevOps, and DataOps. Working knowledge of Agile, Waterfall, DevOps, ITIL, and COBIT. Experience with Jira, Confluence, and Git. Excellent communication and collaboration skills, with the ability to partner effectively with architects, business analysts, DBAs, QA, and cross-functional technology teams. Contract/Hybrid Toronto

What you’ll do

Responsible for day-to-day reliability and the implementation of Service Level Availability (SLAs) and reliability metrics. Manages high-availability production environments, incident response, and service restoration using AI-powered engineering tools.

Requirements

Requires 6+ years of experience in SRE or related platform engineering roles with strong knowledge of cloud platforms and distributed systems. Proficiency in Python, PowerShell, and various DevOps/DataOps methodologies is essential.

Listed skills

  • CI/CD · Preferred
  • Root Cause Analysis · Preferred
  • Agile · Preferred
  • Python · Preferred

Other relevant skills

Identified from the job description. Confirm important requirements above.

  • SRE
  • Production Engineering
  • Cloud Platforms
  • Incident Response
  • Root Cause Analysis
  • Observability
  • Automation
  • Python
  • PowerShell
  • APIs
  • CI/CD
  • Infrastructure as Code
  • DevOps
  • DataOps
  • Agile
  • ITIL

Job areas

  • Technology
  • Software
  • Engineering
  • Management & Leadership
  • Finance & Accounting

More jobs you can apply to directly

Similar opportunities posted by employers hiring on Jobs.ca, with no external application form.

Browse all Easy Apply jobs