Back to job search
Deep Genomics logo
Deep GenomicsVerified Job Source

Senior MLOps Engineer

  • Toronto, ON
  • Hybrid
  • Posted Apr 9, 2026
  • 1 position

$175,000–$200,000 / year

Opens an external site

Sign in to save this job
Employment type
Full-time
Experience level
Mid-level · 2+ years
Posting language
English
Working hours
40 hours per week

Job summary

You will own and evolve the infrastructure powering ML pipelines, including cloud environments, CI/CD systems, and workflow orchestration. You will collaborate with scientists and engineers to ensure platform reliability, scalability, and reproducibility for AI-driven drug discovery.

Job details

ABOUT US Deep Genomics is at the forefront of using artificial intelligence to transform drug discovery. Our proprietary AI platform decodes the complexity of RNA biology to identify novel drug targets, mechanisms, and therapeutics inaccessible through traditional methods. With expertise spanning machine learning, bioinformatics, data science, engineering, and drug development, our multidisciplinary team in Toronto and Cambridge, MA is revolutionizing how new medicines are created. OPPORTUNITY Join us in building the future of AI-driven drug discovery as a Senior MLOps Engineer. You will own and evolve the infrastructure that powers our ML pipelines – from cloud environments and CI/CD systems to workflow orchestration and model deployment. You will work closely with ML scientists, bioinformaticians, and software engineers to keep our platform reliable, reproducible, and scalable. IDEAL CANDIDATE You are someone who enjoys keeping the infrastructure running smoothly so that scientists can focus on their research. You are comfortable working across cloud platforms, CI/CD systems, containers, and GPUs – and you take pride in making these systems reliable and easy for others to use. You have 4+ years of experience in production infrastructure or MLOps, you write solid Python, and you are curious about the ML and scientific workflows your work supports. Above all, you are a collaborative, kind team member who communicates clearly, adapts to evolving needs, and is happy to help colleagues grow their own infrastructure skills along the way. If this sounds like you, we would love to hear from you. \n Key Responsibilities * Maintain and improve cloud infrastructure (GCP) using Infrastructure-as-Code tools (Terraform). * Manage IAM, RBAC, and permission policies across cloud environments. * Own and evolve CI/CD pipelines (CircleCI, GitHub Actions) and ensure best practices are followed across the engineering and ML teams. * Administer and support workflow orchestration platforms (e.g., Seqera/Nextflow, Argo, Kubeflow). * Operate and configure ML experiment tracking and registry tooling (e.g., W&B, MLflow). * Build and maintain containerized environments (Docker) and manage Kubernetes clusters. * Manage GPU resources – provisioning, scheduling, and debugging hardware and driver issues. * Write and maintain Python tooling, scripts, and integrations that support ML infrastructure. * Help deploy ML models to production environments and monitor their performance. Basic Qualifications * 4+ years of experience operating production infrastructure. * Proficiency with cloud platforms (GCP preferred; AWS/Azure acceptable) and Infrastructure-as-Code (Terraform). * Extensive Hands-on experience with Kubernetes and containerization (Docker). * Solid background in CI/CD systems (CircleCI, GitHub Actions, or similar). * Experience managing GPU compute (provisioning, debugging, driver management). * Familiarity with Python package and environment management (e.g., pip, conda, pixi). * Strong Python programming skills. * Self-motivated problem solver with excellent communication skills. Preferred Qualifications * Understanding of ML frameworks (e.g., PyTorch, PyTorch Lightning), ML workflows (training, inference, evaluation), and the model lifecycle. * Familiarity with MLOps tooling (e.g., W&B, Ray, VertexAI) and distributed compute patterns (e.g., DDP, realtime/batch inference, multi-node training). * Familiarity with Kubernetes CRDs and batch/gang schedulers (e.g., Volcano, Kueue). * Experience working with large-scale datasets (storage, versioning, efficient access patterns). * Experience working directly with scientists and researchers in an interdisciplinary setting. * Knowledge of biology and/or machine learning science. * Familiarity with data compliance and governance frameworks (e.g., HIPAA, SOC 2). * Previous startup experience. What We Offer * A collaborative and innovative environment at the frontier of computational biology, machine learning, and drug discovery. * Highly competitive compensation, including meaningful stock ownership. * Comprehensive benefits - including health, vision, and dental coverage for employees and families, employee and family assistance program. * Flexible work environment - including flexible hours, extended long weekends, holiday shutdown, unlimited personal days. * Maternity and parental leave top-up coverage, as well as new parent paid time off. * Focus on learning and growth for all employees - learning and development budget & lunch and learns. * Facilities located in the heart of Toronto - the epicenter of machine learning and AI research and development, and in Kendall Square, Cambridge, Mass. - a global center of biotechnology and life sciences. \n Deep Genomics encourages applications from all backgrounds who seek the opportunity to build the world's leading AI-driven genetic medicine company. If you have a disability or special need, accommodation is available on request for candidates taking part in all aspects of the selection process. *This posting reflects a current vacancy. We offer competitive compensation aligned with local market benchmarks. The salary range for this role is $175,000 - $200,000, and reflects Canada-based roles; compensation may differ for U.S.-based candidates.

What you’ll do

You will own and evolve the infrastructure powering ML pipelines, including cloud environments, CI/CD systems, and workflow orchestration. You will collaborate with scientists and engineers to ensure platform reliability, scalability, and reproducibility for AI-driven drug discovery.

Requirements

The role requires 4+ years of experience in production infrastructure or MLOps with strong Python programming skills. Proficiency in cloud platforms, Kubernetes, containerization, and Infrastructure-as-Code tools like Terraform is essential.

Benefits

  • Health coverage
  • Vision coverage
  • Dental coverage
  • Stock ownership
  • Employee and family assistance program
  • Flexible hours
  • Extended long weekends
  • Holiday shutdown
  • Unlimited personal days
  • Maternity and parental leave top-up
  • New parent paid time off
  • Learning and development budget
  • Lunch and learns

Listed skills

  • Kubernetes · Preferred
  • Budget · Preferred
  • Production · Preferred
  • collaborative · Preferred
  • CI/CD · Preferred
  • Reliability · Preferred
  • Machine learning · Preferred
  • Development · Preferred
  • batch · Preferred
  • Flexible · Preferred
  • Terraform · Preferred
  • Python · Preferred

Other relevant skills

Identified from the job description. Confirm important requirements above.

  • Python
  • MLOps
  • Kubernetes
  • Terraform
  • GCP
  • CI/CD
  • Docker
  • GPU management
  • Workflow orchestration
  • Nextflow
  • Argo
  • Kubeflow
  • MLflow
  • Infrastructure-as-Code
  • PyTorch
  • Distributed compute

Job areas

  • Technology
  • Engineering
  • Science & Research
  • Software
  • Data & Analytics

More jobs from Deep Genomics

See all jobs from Deep Genomics