Cloud Engineer ( Azure Platform Engineer)- Remote Ontario
Top Benefits
About the role
Working for a company like Smile Digital Health means supporting our mandate for #BetterGlobalHealth. We strive towards this goal every day, and the results can be seen in the impact of our innovative health data platform and data management solutions, which are used in over 20 countries. We were #19 on Deloitte's Technology Fast 50 Ranking for 2024! Smile Digital Health makes it easy for healthcare stakeholders to collect and exchange data with our leading FHIR-based data liberation platform. At its heart, the Smile platform enables people and organizations to better manage healthcare data. We help generate and liberate structured healthcare data to ensure effective delivery across care teams and health systems bringing #BetterGlobalHealth to patients everyday!
Apply today and find plenty of reasons to SMILE!
The Azure Platform Engineer is responsible for designing, deploying, and operating production-grade Azure Kubernetes Service (AKS) clusters, with a strong focus on cluster internals, networking, security hardening, and reliability. This role owns the platform end-to-end from cluster architecture and CI/CD through to performance tuning, observability, and incident response for critical systems. Acting as the SME for Kubernetes deployments, blockers, and production issues. The successful candidate works closely with application, security, and platform teams to keep AKS environments reliable, observable, and production-ready. \n
Responsibilities
Act as the subject-matter expert (SME) for Kubernetes deployments, troubleshooting, and production issues across all environments. Design and deploy Azure Kubernetes Service (AKS) clusters with private cluster configurations, managed identities, and RBAC. Own the health, scaling, and lifecycle management of production AKS clusters, including upgrades, node pool management, autoscaling, and capacity planning. Configure AKS networking, including Azure CNI, internal load balancers, and ingress controllers (NGINX, Traefik). Design and maintain integrations between AKS with Azure Container Registry (ACR), Key Vault via CSI driver, Azure Monitor for containers, Azure SQL, Kafka/Event Hubs, Azure Storage, and other client-facing dependencies (DNS resolution, firewall rules, private endpoints, and service connectivity). Manage containerized application deployments using Docker and Helm; maintain reusable chart and templating standards, namespaces, resource quotas, and Azure Policy for AKS. Harden AKS environments through policy enforcement, network policies, and image scanning. Own container and cluster vulnerability management: scanning, triage, prioritization, and remediation coordination with engineering teams. Contribute to Terraform-based infrastructure as code for provisioning and managing Azure resources. Support Azure DevOps (or equivalent) CI/CD pipelines, including GitOps workflows (Flux/ArgoCD), to reduce deployment risk and improve release velocity. Design, implement, and manage observability across the Grafana stack (Prometheus, Loki, Tempo) and Azure-native tooling (Azure Monitor, Log Analytics Workspace, Application Insights) for critical systems; define SLIs/SLOs and tune alerting to reduce noise. Collaborate with performance engineering and application teams to identify, diagnose, and resolve performance bottlenecks spanning the AKS platform and its dependent services (database, messaging, network) to right-size node pools and workloads based on observed performance and utilization trends. Contributing to Disaster Recovery and Business Continuity Planning (DR/BCP) procedures for AKS-hosted workloads, including cross-team failover drill participation and RTO/RPO validation. Provide escalation support for production Kubernetes and infrastructure incidents; participate in on-call rotation, lead root-cause analysis, and drive preventative follow-up actions. Document runbooks and post-incident reviews, and operational knowledge; maintain a living knowledge base to reduce tribal knowledge.
5+ years of hands-on Kubernetes experience, including at least 2 years running AKS in production. Strong knowledge of Kubernetes internals scheduling, networking, storage, RBAC. Proficiency in Azure CNI networking and AKS private cluster configuration. Hands-on experience integrating AKS with Azure PaaS services (ACR, AKV, Azure SQL, Kafka and Managed Identities) and troubleshooting network-layer dependencies (DNS, firewall, private endpoints). Experience with Helm and GitOps workflows (Flux/ArgoCD). Working knowledge of Terraform for infrastructure as code. Working knowledge of Azure DevOps or similar CI/CD tooling. Hands-on experience implementing and operating observability platforms: Grafana stack (Prometheus, Loki, Tempo) and Azure-native monitoring (Azure Monitor, Log Analytics, Application Insights); experience defining SLIs/SLOs for production systems. Practical experience with container/cluster vulnerability management and remediation workflows. Scripting skills in Bash, Python. Certified Kubernetes Administrator (CKA), Or CKAD certification and Microsoft Certified: Azure Administrator (AZ-104) required. Excellent written and verbal communication skills; able to convey technical issues clearly to both technical and non-technical stakeholders.
\n $115,000 - $130,000 a year \n Smile discloses that artificial intelligence (AI) may be used in portions of the recruitment and selection process, such as resume screening or application assessment. All hiring decisions are ultimately made by qualified human decision-makers, and AI tools are used to support — not replace — fair and equitable hiring practices. This position is a new role, created to support Smile’s continued growth and commitment to operational excellence.
Some of the benefits we offer
- Remote Work Environment
- Flexible Time Away From Work Policy including PTO, Personal and Sick Days
- Competitive Salary and Health/Medical Benefits
- RRSP/TFSA/401K Employee Contribution
- Life and Disability
- Employee Assistance Program
- FHIR Study Program and Skillsoft Learning
- Super HAPI Fun Club
Smile's core values include respect, inclusion, embracing our differences, and celebrating shared values because our people are the foundation of our success. We are big on creating a sense of belonging and empowering each other to bring our authentic selves to work. We are dedicated to fostering a workplace that values diversity, equity, and inclusion. We welcome and encourage candidates of all backgrounds to apply. Candidates are encouraged to inform us if they wish to discuss or require accommodations during interviews or while working at Smile.
Not the right fit? Search for Cloud Engineer jobs in Toronto, Ontario, Canada
About Smile Digital Health
Smile Digital Health (doing business as Smile CDR Inc.) specializes in delivering fast, secure, compliant data infrastructures as a service to enable and empower interconnectivity for data-intensive sectors such as healthcare. We are a solutions platform helping organizations like governments, researchers, health systems, healthcare providers, and app developers build connected health solutions and products by leveraging our core expertise in health data and HL7 FHIR.
Our flagship product, Smile CDR, is the world’s first FHIR-based clinical data repository (CDR) as a service. Smile CDR is a high performance and secure solution built on the principles of Privacy by Design. Flexibility is a hallmark of the service as it can be hosted in the cloud or on-premises depending on your needs. The design was strongly influenced by real-world experience in both the jurisdictional and organizational setting. Smile CDR provides a rich set of features and capabilities including:
-Multiple FHIR versions -FHIR Profiles -Full text indexing of clinical records -Type ahead search functionality -FHIR, HL7v2 and custom ETL for data input -Federated identity and identity provider functionality -International character locale support -Terminology services -Auditing -Smart on FHIR support -Rapid deployment -Extensive administrative tooling
Leveraging more than 40 years of experience in building enterprise-class systems, our team includes experts in building integrated healthcare systems including FHIR-based solutions such as HAPI, e-prescriptions and other e-health solutions. We were motivated by the need to improve upon existing options for sharing health data within and across organizations. Smile CDR is the maintainer of HAPI FHIR, the prevailing open source reference implementation of FHIR worldwide
Similar Jobs
Cloud Engineer ( Azure Platform Engineer)- Remote Ontario
Top Benefits
About the role
Working for a company like Smile Digital Health means supporting our mandate for #BetterGlobalHealth. We strive towards this goal every day, and the results can be seen in the impact of our innovative health data platform and data management solutions, which are used in over 20 countries. We were #19 on Deloitte's Technology Fast 50 Ranking for 2024! Smile Digital Health makes it easy for healthcare stakeholders to collect and exchange data with our leading FHIR-based data liberation platform. At its heart, the Smile platform enables people and organizations to better manage healthcare data. We help generate and liberate structured healthcare data to ensure effective delivery across care teams and health systems bringing #BetterGlobalHealth to patients everyday!
Apply today and find plenty of reasons to SMILE!
The Azure Platform Engineer is responsible for designing, deploying, and operating production-grade Azure Kubernetes Service (AKS) clusters, with a strong focus on cluster internals, networking, security hardening, and reliability. This role owns the platform end-to-end from cluster architecture and CI/CD through to performance tuning, observability, and incident response for critical systems. Acting as the SME for Kubernetes deployments, blockers, and production issues. The successful candidate works closely with application, security, and platform teams to keep AKS environments reliable, observable, and production-ready. \n
Responsibilities
Act as the subject-matter expert (SME) for Kubernetes deployments, troubleshooting, and production issues across all environments. Design and deploy Azure Kubernetes Service (AKS) clusters with private cluster configurations, managed identities, and RBAC. Own the health, scaling, and lifecycle management of production AKS clusters, including upgrades, node pool management, autoscaling, and capacity planning. Configure AKS networking, including Azure CNI, internal load balancers, and ingress controllers (NGINX, Traefik). Design and maintain integrations between AKS with Azure Container Registry (ACR), Key Vault via CSI driver, Azure Monitor for containers, Azure SQL, Kafka/Event Hubs, Azure Storage, and other client-facing dependencies (DNS resolution, firewall rules, private endpoints, and service connectivity). Manage containerized application deployments using Docker and Helm; maintain reusable chart and templating standards, namespaces, resource quotas, and Azure Policy for AKS. Harden AKS environments through policy enforcement, network policies, and image scanning. Own container and cluster vulnerability management: scanning, triage, prioritization, and remediation coordination with engineering teams. Contribute to Terraform-based infrastructure as code for provisioning and managing Azure resources. Support Azure DevOps (or equivalent) CI/CD pipelines, including GitOps workflows (Flux/ArgoCD), to reduce deployment risk and improve release velocity. Design, implement, and manage observability across the Grafana stack (Prometheus, Loki, Tempo) and Azure-native tooling (Azure Monitor, Log Analytics Workspace, Application Insights) for critical systems; define SLIs/SLOs and tune alerting to reduce noise. Collaborate with performance engineering and application teams to identify, diagnose, and resolve performance bottlenecks spanning the AKS platform and its dependent services (database, messaging, network) to right-size node pools and workloads based on observed performance and utilization trends. Contributing to Disaster Recovery and Business Continuity Planning (DR/BCP) procedures for AKS-hosted workloads, including cross-team failover drill participation and RTO/RPO validation. Provide escalation support for production Kubernetes and infrastructure incidents; participate in on-call rotation, lead root-cause analysis, and drive preventative follow-up actions. Document runbooks and post-incident reviews, and operational knowledge; maintain a living knowledge base to reduce tribal knowledge.
5+ years of hands-on Kubernetes experience, including at least 2 years running AKS in production. Strong knowledge of Kubernetes internals scheduling, networking, storage, RBAC. Proficiency in Azure CNI networking and AKS private cluster configuration. Hands-on experience integrating AKS with Azure PaaS services (ACR, AKV, Azure SQL, Kafka and Managed Identities) and troubleshooting network-layer dependencies (DNS, firewall, private endpoints). Experience with Helm and GitOps workflows (Flux/ArgoCD). Working knowledge of Terraform for infrastructure as code. Working knowledge of Azure DevOps or similar CI/CD tooling. Hands-on experience implementing and operating observability platforms: Grafana stack (Prometheus, Loki, Tempo) and Azure-native monitoring (Azure Monitor, Log Analytics, Application Insights); experience defining SLIs/SLOs for production systems. Practical experience with container/cluster vulnerability management and remediation workflows. Scripting skills in Bash, Python. Certified Kubernetes Administrator (CKA), Or CKAD certification and Microsoft Certified: Azure Administrator (AZ-104) required. Excellent written and verbal communication skills; able to convey technical issues clearly to both technical and non-technical stakeholders.
\n $115,000 - $130,000 a year \n Smile discloses that artificial intelligence (AI) may be used in portions of the recruitment and selection process, such as resume screening or application assessment. All hiring decisions are ultimately made by qualified human decision-makers, and AI tools are used to support — not replace — fair and equitable hiring practices. This position is a new role, created to support Smile’s continued growth and commitment to operational excellence.
Some of the benefits we offer
- Remote Work Environment
- Flexible Time Away From Work Policy including PTO, Personal and Sick Days
- Competitive Salary and Health/Medical Benefits
- RRSP/TFSA/401K Employee Contribution
- Life and Disability
- Employee Assistance Program
- FHIR Study Program and Skillsoft Learning
- Super HAPI Fun Club
Smile's core values include respect, inclusion, embracing our differences, and celebrating shared values because our people are the foundation of our success. We are big on creating a sense of belonging and empowering each other to bring our authentic selves to work. We are dedicated to fostering a workplace that values diversity, equity, and inclusion. We welcome and encourage candidates of all backgrounds to apply. Candidates are encouraged to inform us if they wish to discuss or require accommodations during interviews or while working at Smile.
Not the right fit? Search for Cloud Engineer jobs in Toronto, Ontario, Canada
About Smile Digital Health
Smile Digital Health (doing business as Smile CDR Inc.) specializes in delivering fast, secure, compliant data infrastructures as a service to enable and empower interconnectivity for data-intensive sectors such as healthcare. We are a solutions platform helping organizations like governments, researchers, health systems, healthcare providers, and app developers build connected health solutions and products by leveraging our core expertise in health data and HL7 FHIR.
Our flagship product, Smile CDR, is the world’s first FHIR-based clinical data repository (CDR) as a service. Smile CDR is a high performance and secure solution built on the principles of Privacy by Design. Flexibility is a hallmark of the service as it can be hosted in the cloud or on-premises depending on your needs. The design was strongly influenced by real-world experience in both the jurisdictional and organizational setting. Smile CDR provides a rich set of features and capabilities including:
-Multiple FHIR versions -FHIR Profiles -Full text indexing of clinical records -Type ahead search functionality -FHIR, HL7v2 and custom ETL for data input -Federated identity and identity provider functionality -International character locale support -Terminology services -Auditing -Smart on FHIR support -Rapid deployment -Extensive administrative tooling
Leveraging more than 40 years of experience in building enterprise-class systems, our team includes experts in building integrated healthcare systems including FHIR-based solutions such as HAPI, e-prescriptions and other e-health solutions. We were motivated by the need to improve upon existing options for sharing health data within and across organizations. Smile CDR is the maintainer of HAPI FHIR, the prevailing open source reference implementation of FHIR worldwide