Senior Infrastructure Engineer (Reign)
- Canada, United States
- Remote
- Posted Sep 14, 2026
- 1 position
Opens an external site
- Employment type
- Full-time
- Experience level
- Senior · 5+ years
- Posting language
- English
- Working hours
- 40 hours per week
Job summary
Design and maintain modular infrastructure definitions that allow clients to deploy and operate the platform in diverse environments, including public clouds and air-gapped on-prem clusters. Manage the full lifecycle of infrastructure and deployment changes to ensure predictability, high availability, and robust security across all client environments.
Job details
iTmethods · Reign engineering · Remote (Canada) · Full time · Reports to VP Engineering We use artificial intelligence tools to help screen and assess applications. A human makes advancement and hiring decisions. Reign is an AI governance platform for regulated enterprises. You take it from a deployment we run ourselves to one a client can install, operate, and keep running in their own environment. Our clients run the platform in their own environment, on their own cloud, under their own controls. Some run managed Kubernetes on a public cloud. Some run on-prem clusters with no internet access. You make the platform deployable in all of those places, at scale, without help from our engineers. The infrastructure is built from modular definitions that each client composes for their environment. You design those modules, maintain them as the platform changes, and manage how infrastructure and deployment changes reach each client without breaking what already runs. What you will do Design the reference deployment and make it portable across cloud providers and on-prem environments. Design the infrastructure as modules that clients compose for their environment, and maintain the definitions and modules as the platform changes. Manage infrastructure and deployment changes from proposal through review, rollout, and rollback, so each client environment stays predictable. Build the install, upgrade, and rollback path that a client operator can run. Make every layer of the platform highly available. Define the failure modes and prove each one with a test. Evolve the deployment and its maintainability as the client base grows, so each new client is cheaper to onboard and operate than the last. Own the environments and the pipelines that deploy them. Own the observability that ships with the platform and works with the client's own tooling. Own supply-chain security, including signed artifacts, vulnerability gates, and offline installs. Review architecture with the team, and bring the deployment and operations view in early so every design is one a client can run. What you bring (required) You have designed systems that other people operate, and you have lived with the consequences. You name failure modes up front and prove recovery with a test. You keep designs simple enough to explain why each component exists. You treat install, upgrade, rollback, and security as part of the design, not an afterthought. You have run production workloads on Kubernetes across more than one cloud or on-prem environment. You can defend a design to engineers who disagree with it. You recognize code and infrastructure patterns, and you understand how architecture drifts over time and what it takes to correct it. Nice to have You have delivered software into air-gapped or sovereign-cloud environments. You have packaged software for clients to install and operate themselves. You have shipped into a regulated industry. You have run AI coding agents at scale. Companion role: Senior Software Engineer, Reign.
What you’ll do
Design and maintain modular infrastructure definitions that allow clients to deploy and operate the platform in diverse environments, including public clouds and air-gapped on-prem clusters. Manage the full lifecycle of infrastructure and deployment changes to ensure predictability, high availability, and robust security across all client environments.
Requirements
Requires extensive experience running production Kubernetes workloads across multiple cloud or on-prem environments and a proven ability to design systems that are maintainable by others. Candidates must demonstrate a strong focus on failure mode analysis, security, and the ability to defend architectural decisions to technical stakeholders.
Listed skills
- Kubernetes · Preferred
Other relevant skills
Identified from the job description. Confirm important requirements above.
- Kubernetes
- Infrastructure as Code
- Cloud Computing
- On-premise Infrastructure
- System Architecture
- Deployment Pipelines
- Observability
- Supply-chain Security
- Automation
- Failure Mode Analysis
- Software Packaging
- Regulated Industry Compliance
- Air-gapped Environments
- AI Governance
- Pipelines
- Sovereign Cloud (Data Security Framework)
- Supply Chain Security
- Artificial Intelligence
- Software Development
- Failure Causes
- Governance
- Operations
- Public Cloud
- Tooling
- On Prem
Job areas
- Technology
- Engineering
- Software
- Infrastructure Engineer
- Systems Engineer
- Software Developers
- Computer Systems Engineers/Architects
- Computer Occupations, All Other
More jobs you can apply to directly
Similar opportunities posted by employers hiring on Jobs.ca, with no external application form.
Trista
Residential Home Care Services Manager (NOC 60040)
SponsoredDirect employerEasy Apply- On-site
- Posted Sep 11, 2026
Desjardins
Senior Litigation Advisor -
SponsoredDirect employerEasy Apply- Hybrid
- Posted Sep 9, 2026
Bédard Ressources Humaines
Adjoint(e) administratif(ve) à la direction #381
SponsoredDirect employerEasy Apply- On-site
- Posted Sep 9, 2026
