RPE - System Reliability Engineering Specialist (Hybrid)
- Montréal, QC
- Hybrid
- Posted Aug 31, 2026
- 1 position
Opens an external site
- Employment type
- Full-time
- Experience level
- Mid-level · 2+ years
- Posting language
- English
- Working hours
- 40 hours per week
Job summary
The role involves supporting the reliability, performance, and operational stability of business-critical applications while driving automation and AI-enabled technologies. Responsibilities include managing major incidents, performing root cause analysis, and maintaining operational tooling to improve efficiency.
Job details
We’re seeking someone to join our team as a Site Reliability Engineering Specialist to support the reliability, performance, and operational stability of business-critical applications, while helping drive automation and the adoption and operation of AI-enabled technologies and modern engineering practices. In the Technology division, we leverage innovation to build the connections and capabilities that power our Firm, enabling our clients and colleagues to redefine markets and shape the future of our communities. This is a Software Production Management & Reliability Engineering II position at Associate, which is part of the job family responsible for overseeing the production environment, ensuring the operational reliability of deployed software, and implementing strategies to optimize performance and minimize downtime. Since 1935, Morgan Stanley is known as a global leader in financial services, always evolving and innovating to better serve our clients and our communities in more than 40 countries around the world. What you'll do in the role: * Work independently to solve ambiguous operational and technical problems while supporting business-critical applications. * Collaborate with development, QA, product management, and business stakeholders to support applications throughout their lifecycle. * Escalate and manage major incidents, perform root cause analysis, and implement corrective actions to improve reliability. * Develop and maintain automation, scripts, operational tooling, documentation, and runbooks to improve efficiency and reduce manual effort. * Perform platform lifecycle activities, including upgrades, patching, archival, cleanup, and ongoing maintenance. * Support the evaluation, deployment, and operation of third-party software, AI-enabled platforms, and emerging technologies. * Partner with global teams to improve support processes, operational effectiveness, and continuous improvement initiatives. What you'll bring to the role: * At least 2 years' relevant experience would generally be expected to find the skills required for this role. * Strong knowledge of Linux/Unix operating systems and production support environments. * Strong scripting and automation skills using Python, Bash, Shell scripting, or similar technologies. * Understanding of networking concepts, protocols, APIs, REST services, system integrations, and troubleshooting methodologies. * Experience with monitoring, observability, and reliability engineering concepts and tools such as OpenTelemetry, Grafana, and Loki. * Experience using modern integrated development environments (IDEs) and version control platforms, including Visual Studio Code and Git. * Familiarity with containerization technologies such as Docker and Kubernetes, and exposure to cloud technologies and automation frameworks. * Basic understanding of Generative AI concepts, including Large Language Models (LLMs), AI agents, prompt engineering, and Retrieval-Augmented Generation (RAG). * Familiarity with AI-assisted development tools such as GitHub Copilot, Microsoft Copilot, or similar productivity platforms, with an interest in applying AI and automation to improve operational efficiency. All our positions are located in Montreal, Quebec. We offer a hybrid work environment, combining remote work and attendance in the office. Knowledge of French and English is required. WHAT YOU CAN EXPECT FROM MORGAN STANLEY: At Morgan Stanley, we raise, manage and allocate capital for our clients – helping them reach their goals. We do it in a way that’s differentiated – and we’ve done that for 90 years. Our values - putting clients first, doing the right thing, leading with exceptional ideas, committing to diversity and inclusion, and giving back - aren’t just beliefs, they guide the decisions we make every day to do what's best for our clients, communities and more than 80,000 employees in 1,200 offices across 42 countries. At Morgan Stanley, you’ll find an opportunity to work alongside the best and the brightest, in an environment where you are supported and empowered. Our teams are relentless collaborators and creative thinkers, fueled by their diverse backgrounds and experiences. We are proud to support our employees and their families at every point along their work-life journey, offering some of the most attractive and comprehensive employee benefits and perks in the industry. There’s also ample opportunity to move about the business for those who show passion and grit in their work. To learn more about our offices across the globe, please copy and paste https://www.morganstanley.com/about-us/global-offices into your browser. Morgan Stanley is an equal opportunity employer committed to building and maintaining a workforce that is diverse in experience and background. Our recruiting efforts reflect our strong commitment to a culture of inclusion, where individuals are hired, developed, and advanced based on their skills and talents. Our workforce reflects a broad cross-section of the global communities in which we operate, bringing a variety of backgrounds, talents, perspectives, and experiences. For more information, please visit: https://www.morganstanley.com/people-opportunities/eeo [https://www.morganstanley.com/people-opportunities/eeo].
What you’ll do
The role involves supporting the reliability, performance, and operational stability of business-critical applications while driving automation and AI-enabled technologies. Responsibilities include managing major incidents, performing root cause analysis, and maintaining operational tooling to improve efficiency.
Requirements
Candidates are expected to have at least 2 years of relevant experience with strong knowledge of Linux/Unix environments and scripting languages like Python or Bash. Familiarity with monitoring tools, containerization, and cloud technologies is required, along with a basic understanding of Generative AI concepts.
Benefits
• Comprehensive employee benefits • Work-life balance support • Professional development opportunities
Listed skills
- Kubernetes · Preferred
- Docker · Preferred
- Linux · Preferred
- prompt engineering · Preferred
- Python · Preferred
Other relevant skills
Identified from the job description. Confirm important requirements above.
- Linux
- Unix
- Python
- Bash
- Shell scripting
- Automation
- Networking
- APIs
- REST services
- OpenTelemetry
- Grafana
- Loki
- Docker
- Kubernetes
- Generative AI
- Prompt engineering
- Microsoft Copilot
- GitHub Copilot
- Large Language Modeling
- Prompt Engineering
- Operational Efficiency
- Generative Artificial Intelligence
- Observability
- Site Reliability Engineering
- Git (Version Control System)
- Retrieval Augmented Generation
- Emerging Technologies
- Bilingual (French/English)
- AI Agents
- Application Programming Interface (API)
- Artificial Intelligence
- Bash (Scripting Language)
- Management
- Containerization
- Version Control
- Continuous Improvement Process
- Development Environment
- Programming Tools
- Financial Services
- Innovation
- Integrated Development Environments
- Python (Programming Language)
- Microsoft Visual Studio
- Product Management
- Reliability Engineering
- RESTful API
- Safety Assurance
- Software Manufacturing
- Tooling
- Troubleshooting (Problem Solving)
Job areas
- Technology
- Software
- Engineering
- Finance & Accounting
- Data & Analytics
- Systems Reliability Engineer
- Systems Engineer
- Software Developers
- Computer Systems Engineers/Architects
- Computer Occupations, All Other
More jobs you can apply to directly
Similar opportunities posted by employers hiring on Jobs.ca, with no external application form.
Desjardins
Analyste d'affaires système(BSA) Guidewire
Direct employerEasy Apply- Hybrid
- Lévis, QC
- Posted Sep 21, 2026
Desjardins
Administrateur ou administratrice de plateforme TI - Senior
Direct employerEasy Apply- Hybrid
- Montréal, QC
- Posted Sep 17, 2026
Bédard Ressources Humaines
Responsable d’expédition
SponsoredDirect employerEasy Apply- On-site
- Mascouche, QC
- Posted Sep 16, 2026
