Customer Reliability Engineer
Lead troubleshooting investigations for complex customer problems and act as the primary technical point of contact. Own core reliability workstreams including incident response, backup/DR validation, and security patching.
- Remote
- Canada
- Posted Jul 10, 2026
- 1 position
Job summary
OpsGuru, a Carbon60 Company, is a global engineering and consulting group that helps organizations accelerate digital transformation through cutting-edge technology, deep engineering expertise, and outcome-driven solutions. OpsGuru's value to our customers centres around our ability to provide deep technical guidance based on their business needs. We achieve this by assigning small, virtual teams of highly skilled individuals to each client. Minimum Qualifications: +7 years of experience in Information technology, preference for cloud engineering, site reliability engineering, infrastructure or platform services operations. +4 hands-on experience with AWS Cloud infrastructure and ecosystem (EC2, RDS, S3, VPC, IAM, KMS, Backup, CloudTrail, Route53, ECS/EKS, etc.) Proficiency with Network and Security troubleshooting (Routing, DNS, Security Groups, ACL’s, WAF, VPN/Bastion access). Proficiency in at least one scripting language such as Ruby, Python, Bash or Powershell Proficiency in at least one of the following Infrastructure as Code frameworks: Terraform, CDK, CloudFormation Previous Production On call 24/7 experience, ideally with an incident management/paging tool (e.g., PagerDuty) - MUST Previous Windows/Linux server administration and patching experience Familiarity with at least one of the following monitoring platform: CloudWatch, Grafana, New Relic, Datadog, or Prometheus Experience working with Claude Code, Codex or similar AI driven development environment Nice to haves: Experience with Docker and Kubernetes (EKS/ECS), including container image patching/vulnerability remediation Experience with CI/CD tooling: Git, Jenkins, CircleCI, GitHub/GitLab Experience with data stores such as: MySQL, MSSQL, PostgreSQL Experience with cloud governance/compliance or CSPM tooling (e.g., Lacework, AWS Config, Trusted Advisor) Experience with ITSM/ticketing platforms (e.g., ServiceNow, Jira) AWS Associate level certification (Solutions Architect, SysOps, Devops) Exposure to Azure (AKS, VNet, Azure Backup) — for engagements spanning multi-cloud accounts Your focus: Lead troubleshooting investigations to bring quicker issue resolution to complex problems impacting our customers. Act as the customer's technical single point of contact, establishing credibility and building impactful relationships Own core reliability workstreams: incident response, backup/DR validation, security patching, secure access management, and audit log monitoring Enabling the Operations team to deliver excellent customer service through great documentation and knowledge transfer Work with a wide variety customers, industries, and bleeding-edge technologies Develop skills and contributions through educational and personal growth opportunities What's in it for you: Compensation & Perks Competitive compensation package Retirement Savings Matching Program (RRSP) Partnership with Perkopolis Discounts Flexibility & Time Off Remote first work environment Flexible work hours & location Paid parental leave options Health & Wellness Employer paid health & dental premiums GreenShield+ Counselling Mental Health $500 in Health Care Spending Account annually At OpsGuru, a Carbon60 Company, we encourage employees to bring their whole, authentic selves to work. By sharing and embracing unique backgrounds, experiences, and perspectives, we learn from each other, innovate, and create a dynamic environment where we can be and achieve our best. We're dedicated to ensuring each member feels a sense of belonging, safety, and respect. At Carbon60, your unique voice is heard and embraced, and you meaningfully contribute to decision-making and the organization's growth. Carbon60 is committed to an equitable employee experience, opportunity, and support. If you require accommodations or support during the recruitment process, please email us at careers@opsguru.io We thank all applicants for their interest in this exciting opportunity. Only candidates that meet the qualifications will be contacted for an interview.
What you’ll do
Lead troubleshooting investigations for complex customer problems and act as the primary technical point of contact. Own core reliability workstreams including incident response, backup/DR validation, and security patching.
Requirements
Requires over 7 years of IT experience with at least 4 years of hands-on AWS expertise and proficiency in scripting and Infrastructure as Code. Must have previous 24/7 production on-call experience and server administration skills.
Benefits
• Retirement Savings Matching Program (RRSP) • Perkopolis Discounts • Remote First Work Environment • Flexible Work Hours & Location • Paid Parental Leave Options • Employer Paid Health & Dental Premiums • Counselling • Mental Health Support • Health Care Spending Account
Other relevant skills
Identified from the job description. Confirm important requirements above.
- AWS Cloud Infrastructure
- Network Troubleshooting
- Python
- Terraform
- Incident Management
- Linux Server Administration
- Monitoring Platforms
- Docker
- Kubernetes
- CI/CD
- MySQL
- Cloud Governance
- ITSM
- Azure
- SRE
- Cloud Engineering
Job areas
- Technology
- Engineering
- Consulting
- Software
- Customer Service & Support
Additional details
- Minimum experience
- 5+ years
- Posting language
- English
- Working hours
- 40 hours per week
