Jobs.ca
Jobs.ca
Language
Jobgether logo

Site Reliability Specialist, IT Operations

Jobgether1 day ago
Hybrid
Canada
Mid Level
Full-Time

Top Benefits

Hybrid work environment
Paid vacation
Personal days

About the role

This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Site Reliability Specialist, IT Operations based in Canada. Join a highly technical IT Operations team where you'll help build and maintain reliable, scalable, and resilient production platforms. In this role, you will apply Site Reliability Engineering (SRE) principles to improve system performance, automate operational processes, and enhance service availability across critical environments. Working closely with infrastructure, development, DevOps, security, and product teams, you will play a key role in driving operational excellence through engineering and automation. This position is ideal for professionals who enjoy solving complex technical challenges, reducing manual effort, and improving observability across modern cloud-based environments. You'll have the opportunity to work with cutting-edge technologies while contributing to the continuous evolution of mission-critical services. \n

Accountabilities

As a Site Reliability Specialist, you will be responsible for improving the reliability, performance, and scalability of production systems while reducing operational complexity through automation and engineering best practices. Apply Site Reliability Engineering (SRE) principles to improve the availability, resilience, and performance of production platforms and hosted services. Design, develop, and maintain automation scripts, workflows, and operational tools to reduce manual processes and improve efficiency. Build and support production environments while ensuring compliance with operational procedures, security standards, and reliability best practices. Define and support Service Level Objectives (SLOs), Service Level Indicators (SLIs), and operational reliability standards. Investigate incidents, perform root cause analysis, and implement both immediate fixes and long-term reliability improvements. Enhance monitoring, logging, alerting, telemetry, and overall observability to proactively identify and resolve system issues. Contribute to Infrastructure-as-Code, Configuration-as-Code, Observability-as-Code, and deployment automation initiatives. Collaborate with development, DevOps, infrastructure, security, architecture, and product teams to ensure operational readiness and continuous improvement. Explore AI-powered automation tools and intelligent workflows to optimize operations and accelerate incident response. Create and maintain technical documentation, including SOPs, runbooks, troubleshooting guides, and operational knowledge resources. Manage incidents, service requests, and change activities while meeting service level agreements. Participate in scheduled maintenance, platform upgrades, migrations, and an on-call rotation supporting 24/7 operations.

Requirements

The ideal candidate combines strong infrastructure and automation expertise with a software engineering mindset, enabling continuous improvements across production operations and system reliability. College or university degree in Computer Science, Information Technology, Software Development, Engineering, or equivalent practical experience. 3–5 years of experience in systems administration, IT operations, DevOps, infrastructure support, automation, or software development. Experience supporting business-critical production environments and customer-facing platforms. Strong scripting or programming skills using PowerShell, Python, Bash, JavaScript, TypeScript, C#, or similar languages. Proven experience designing, developing, testing, documenting, and maintaining automation solutions. Solid knowledge of Microsoft and/or Linux server administration and production support. Strong troubleshooting, analytical, and root cause analysis skills. Good understanding of distributed systems, networking, system performance, reliability, and service operations. Experience with monitoring, logging, telemetry, observability platforms, and operational analytics. Familiarity with Git, version control, CI/CD pipelines, deployment automation, and code review practices. Experience with Infrastructure-as-Code or automation tools such as Terraform, Ansible, Azure DevOps, GitHub Actions, Docker, or Kubernetes is an asset. Exposure to Azure AI Foundry, Power Automate, AI agents, or similar automation technologies is considered a plus. Knowledge of cloud platforms, virtualization, backup strategies, and high-availability environments is advantageous. Strong communication and collaboration skills with the ability to work effectively across technical and non-technical teams. Excellent English communication skills; French proficiency is an asset. Industry certifications related to Azure, Linux, Kubernetes, DevOps, or observability platforms are considered beneficial. Availability to participate in a 24/7 rotational on-call schedule.

Benefits

Competitive annual salary ranging from $71,330 to $101,900 CAD, based on experience and qualifications. Hybrid work environment with flexible working arrangements. Modern collaborative office spaces and access to advanced technologies. Paid vacation and personal days starting from day one. Comprehensive employee benefits and savings programs. Professional development opportunities, ongoing training, and career growth support. Open communication culture with regular feedback and development planning. Employee referral bonus program. Diverse, collaborative, and inclusive workplace culture. Engaging virtual and in-person social events throughout the year. Opportunity to work with modern cloud, automation, AI, and Site Reliability Engineering technologies.

\n How Jobgether works: We use an AI-powered matching process to ensure your application is reviewed quickly, objectively, and fairly against the role's core requirements. Our system identifies the top-fitting candidates, and this shortlist is then shared directly with the hiring company. The final decision and next steps (interviews, assessments) are managed by their internal team. We appreciate your interest and wish you the best! Why Apply Through Jobgether? Data Privacy Notice: By submitting your application, you acknowledge that Jobgether will process your personal data to evaluate your candidacy and share relevant information with the hiring employer. This processing is based on legitimate interest and pre-contractual measures under applicable data protection laws (including GDPR). You may exercise your rights (access, rectification, erasure, objection) at any time. #LI-CL1

About Jobgether

Internet Marketplace Platforms
11-50
Founded in 2019

Your future of work, like you've always dreamt it, is now possible with Jobgether !

The Covid crisis has accelerated its revolution but work, as we knew it, doesn't exist anymore. Tomorrow, jobs will be hybrid, remote and asynchronous. Flexibility will be the norm.

Jobgether helps you find your next remote job, wherever you are.

Similar Jobs