Production Support Engineer
The role involves managing system reliability through proactive monitoring, incident response, and automation of repetitive tasks. Additionally, the engineer will manage infrastructure, plan for capacity, and collaborate with developers to ensure seamless software releases.
- On-site
- Toronto, ON
- Posted Jul 20, 2026
- Apply by Aug 19, 2026
- 1 position
Job summary
Choosing Capgemini means choosing a company where you will be empowered to shape your career in the way you’d like, where you’ll be supported and inspired by a collaborative community of colleagues around the world, and where you’ll be able to reimagine what’s possible. Join us and help the world’s leading organizations unlock the value of technology and build a more sustainable, more inclusive world. Role Summary Capgemini is looking for a Production Support Engineer (SRE) to work for the Commercial Line of Business. An ideal candidate should be the one that has experience as outlined below:- Job Description Monitoring and Alerting: Implement and maintain monitoring systems to proactively identify potential issues and alert engineers to problems before they impact users. Incident Response: Respond to incidents and outages, diagnose problems, and implement solutions to minimize downtime and restore service. Automation: Automate repetitive tasks and processes to improve efficiency and reduce manual effort. Infrastructure Management: Manage and maintain the underlying infrastructure, including servers, networks, and cloud resources. Capacity Planning: Plan for future capacity needs to ensure systems can handle anticipated workloads. Release Engineering: Develop and maintain processes for deploying software updates and releases. Collaboration: Work closely with developers, operations teams, and other stakeholders to ensure system reliability and availability. Documentation: Maintain clear and concise documentation of systems, processes, and procedures. Continuous Improvement: Identify areas for improvement and implement changes to enhance system reliability and performance. Skills and Qualifications ● 8+ Years experience in production support handling Prod incidents. ● Excellent knowledge of OCP and windows environment.. ● Monitoring tools (Dynatrace) ● Operating System (Windows, Linux) ● Scripting (Shell Scripting, Python, Power Shell) ● Database (SQL database management, MongoDb) ● Container Services (Kubernetes) ● Disaster Recovery Planning and execution.
What you’ll do
The role involves managing system reliability through proactive monitoring, incident response, and automation of repetitive tasks. Additionally, the engineer will manage infrastructure, plan for capacity, and collaborate with developers to ensure seamless software releases.
Requirements
Candidates must have over 8 years of experience in production support and handling production incidents. Proficiency in OCP, Windows, Linux, monitoring tools like Dynatrace, and scripting languages such as Python or Shell is required.
Other relevant skills
Identified from the job description. Confirm important requirements above.
- Production Support
- SRE
- Monitoring
- Incident Response
- Automation
- Infrastructure Management
- Capacity Planning
- Release Engineering
- OCP
- Windows
- Dynatrace
- Linux
- Shell Scripting
- Python
- PowerShell
- Kubernetes
Job areas
- Technology
- Software
- Engineering
- Data & Analytics
- Finance & Accounting
Additional details
- Minimum experience
- 8+ years
- Apply by
- Aug 19, 2026
- Posting language
- English
- Working hours
- 40 hours per week
- Seniority
- Mid-Senior level
