Storage Site Reliability Engineer (GPFS / Storage Ops) | Ingénieur(e) SRE Stockage (GPFS / Opérations de stockage)
- Montréal, QC
- Hybrid
- Posted Sep 12, 2026
- 1 position
$100,000–$115,000 / year
Opens an external site
- Employment type
- Full-time
- Experience level
- Senior · 5+ years
- Apply by
- Oct 11, 2026
- Posting language
- English
- Working hours
- 40 hours per week
- Office presence
- 3 days per week
- Seniority
- Associate
- Application method
- Direct apply is available
Job summary
Manage and support large-scale enterprise storage environments, specifically focusing on GPFS and NAS technologies. Automate operational processes using Python to improve platform reliability and performance.
Job details
We are looking for a Senior Storage Site Reliability Engineer (SRE) to join a high-performing infrastructure team responsible for designing, automating, and maintaining large-scale enterprise storage platforms. This role is ideal for someone with strong Linux administration, storage technologies, and automation expertise who enjoys solving complex reliability and performance challenges. Hybrid work - 3 days in office in Montreal What You'll Do ✅ Manage and support enterprise storage environments, including GPFS and NAS technologies ✅ Automate operational processes and infrastructure tasks using Python ✅ Monitor, troubleshoot, and improve storage platform reliability and performance ✅ Support Linux-based environments and distributed storage systems ✅ Collaborate with network, cloud, and infrastructure teams to resolve complex issues ✅ Drive continuous improvement through automation, tooling, and operational excellence What We're Looking For ✔ Strong experience with Linux Administration and/or NAS technologies (File Systems, SMB) ✔ Knowledge of Software Defined Storage and Cloud technologies ✔ Experience with storage and Unix protocols including NFS, SMB, Fibre Channel, and iSCSI ✔ Strong Python scripting skills for automation and tooling development ✔ Solid understanding of TCP/IP networking, DNS, CIFS, and related technologies ✔ Experience supporting large-scale infrastructure in production environments ✔ Strong troubleshooting and problem-solving skills Top Skills 🔹 GPFS (IBM Storage Scale) 🔹 Linux Administration 🔹 Python Automation 🔹 NAS Storage (NFS/SMB) 🔹 Software Defined Storage (SDS) 🔹 TCP/IP Networking 🔹 Site Reliability Engineering (SRE) Nous recherchons un(e) Ingénieur(e) principal(e) SRE Stockage pour rejoindre une équipe infrastructure performante responsable de la gestion, de l'automatisation et de la fiabilité de plateformes de stockage d'entreprise à grande échelle. Ce rôle convient parfaitement à une personne passionnée par Linux, les technologies de stockage et l'automatisation. En mode hybride - 3 jours en présentiel au bureau à Montréal Responsabilités ✅ Gérer et soutenir les environnements de stockage d'entreprise, incluant GPFS et les technologies NAS ✅ Automatiser les processus opérationnels et les tâches d'infrastructure à l'aide de Python ✅ Surveiller, diagnostiquer et optimiser la performance et la fiabilité des plateformes de stockage ✅ Assurer le support des environnements Linux et des systèmes de stockage distribués ✅ Collaborer avec les équipes réseau, infonuagique et infrastructure ✅ Contribuer à l'amélioration continue grâce à l'automatisation et aux meilleures pratiques SRE Compétences recherchées ✔ Excellente expérience en administration Linux et/ou technologies NAS (systèmes de fichiers, SMB) ✔ Connaissance du stockage défini par logiciel (SDS) et des technologies Cloud ✔ Maîtrise des protocoles de stockage et Unix : NFS, SMB, Fibre Channel, iSCSI ✔ Excellentes compétences en Python pour l'automatisation et le développement d'outils ✔ Bonne compréhension des technologies TCP/IP, DNS, CIFS et des réseaux d'entreprise ✔ Expérience dans des environnements critiques à grande échelle ✔ Fortes capacités d'analyse et de résolution de problèmes Compétences clés 🔹 GPFS (IBM Storage Scale) 🔹 Linux Administration 🔹 Python Automation 🔹 NAS Storage (NFS/SMB) 🔹 Software Defined Storage (SDS) 🔹 TCP/IP Networking 🔹 Site Reliability Engineering (SRE)
What you’ll do
Manage and support large-scale enterprise storage environments, specifically focusing on GPFS and NAS technologies. Automate operational processes using Python to improve platform reliability and performance.
Requirements
Requires strong expertise in Linux administration, storage protocols (NFS, SMB, iSCSI), and Python scripting for automation. Candidates should have experience supporting large-scale production infrastructure and software-defined storage.
Listed skills
- DNS · Preferred
Other relevant skills
Identified from the job description. Confirm important requirements above.
- GPFS
- Linux Administration
- Python Automation
- NAS Storage
- Software Defined Storage
- TCP/IP Networking
- Site Reliability Engineering
- NFS
- SMB
- Fibre Channel
- iSCSI
- DNS
- CIFS
Job areas
- Technology
- Software
- Engineering
- Consulting
More jobs you can apply to directly
Similar opportunities posted by employers hiring on Jobs.ca, with no external application form.
Bédard Ressources Humaines
Responsable d’expédition
SponsoredDirect employerEasy Apply- On-site
- Mascouche, QC
- Posted Sep 16, 2026
Dalfen Ltée
Building Superintendent
SponsoredDirect employerEasy Apply- On-site
- Sainte-Agathe-Des-Monts, QC
- Posted Sep 15, 2026
Sustainable Projects Group
Sustainability consultant/ Team Lead, Energy Consulting NOC 41400
SponsoredDirect employerEasy Apply- Hybrid
- vancouver v5l 4s1, BC
- Posted Sep 17, 2026
