Back to job search
W
WaabiVerified Job Source

Senior / Staff ML Training Optimization Engineer

Build and optimize standardized distributed training frameworks to improve stability and efficiency for research and production. Profile model runtime and memory to eliminate bottlenecks and implement emerging technologies like custom CUDA kernels.

  • Hybrid
  • Toronto, ON
  • Posted May 8, 2026
  • 1 position

Job summary

Waabi, founded by AI visionary Raquel Urtasun, is the leader in Physical AI. With a world-class team, we're unlocking the next era of autonomous transportation with technology that's powering commercial autonomous trucks and robotaxis. Waabi is backed by and partners with world leaders in AI, automotive, logistics, and deep tech. With offices in Toronto, San Francisco, Dallas, and Pittsburgh, Waabi is growing quickly and looking for diverse, innovative and collaborative candidates who want to impact the world in a positive way. To learn more visit: www.waabi.ai You will... - Build standardized distributed training frameworks for research and production, drive our training towards new levels of stability and efficiency. - Comprehensively profile model runtime and memory to pinpoint performance bottlenecks. - Identify and evaluate emerging technologies that can be adopted into Waabi’s training and inference frameworks. Examples include designing new CUDA kernels, quantization-aware training and inference, and compilation/deployment techniques. - Work with researchers and ML engineers on best-practices for optimal resource usage. - Create and improve tooling and dashboards to ensure broad adoption of your work. Qualifications: - MS/PhD or Bachelors degree with a minimum of 4 years of industry experience in Computer Science, Robotics and/or similar technical field(s) of study. - Solid coding proficiency in a variety of coding languages including Python, C++ or Rust. - Experience in deep learning frameworks such as PyTorch or Jax. - Skilled in profiling CPU and GPU code using tools such as PyTorch Profiler and NVIDIA Nsight. - Open-minded and collaborative team player with willingness to help others. - Passionate about self-driving technologies, solving hard problems, and creating innovative solutions. Bonus/nice to have: - Experience in identifying when custom CUDA kernels are needed, and implementing them. - Experience in Bazel in a monorepo environment, and integrating third party packages into dev environments. - Experience with Kubernetes-based training platforms. \n \n The US yearly salary range for this role is: $141,000 - $249,000 in addition to competitive perks & benefits. Waabi’s yearly salary ranges are determined based on several factors in accordance with the Company’s compensation practices. The salary base range is reflective of the minimum and maximum target for new hire salaries for the position across all US locations. Note: The Company provides additional compensation for employees in this role, including equity incentive awards and an annual performance bonus. Perks/Benefits: Waabi provides a competitive benefits package that includes: - Competitive compensation and equity awards. - Health and Wellness benefits encompassing Medical, Dental and Vision coverage. - Unlimited Vacation. - Flexible hours and Work from Home support. - Daily drinks, snacks and catered meals (when in office). - Regularly scheduled team building activities and social events both on-site, off-site & virtually. - World-class facility that includes a gym, games room (ping pong table, video game consoles, board games, etc), multiple collaborative working spaces and a gorgeous patio!(when in office) - As we grow, this list continues to evolve! Waabi is an equal opportunity employer that celebrates diversity and is committed to creating a supportive, inclusive, and accessible environment for all employees. If reasonable accommodation is needed to participate in the job application or interview process please let our recruiting team know.

What you’ll do

Build and optimize standardized distributed training frameworks to improve stability and efficiency for research and production. Profile model runtime and memory to eliminate bottlenecks and implement emerging technologies like custom CUDA kernels.

Requirements

Requires a degree in Computer Science or Robotics with at least 4 years of industry experience and proficiency in Python, C++, or Rust. Candidates must be skilled in deep learning frameworks like PyTorch or Jax and GPU profiling tools.

Benefits

• Competitive Compensation • Equity Awards • Medical Coverage • Dental Coverage • Vision Coverage • Unlimited Vacation • Flexible Hours • Work From Home Support • Daily Drinks • Snacks • Catered Meals • Team Building Activities • Social Events • Gym • Games Room • Collaborative Working Spaces • Patio

Other relevant skills

Identified from the job description. Confirm important requirements above.

  • Distributed Training Frameworks
  • Model Profiling
  • CUDA Kernels
  • Quantization-Aware Training
  • Python
  • C++
  • Rust
  • PyTorch
  • Jax
  • PyTorch Profiler
  • NVIDIA Nsight
  • Bazel
  • Kubernetes
  • Deep Learning
  • Performance Optimization
  • Inference Frameworks

Job areas

  • Software
  • Engineering
  • Technology
  • Data & Analytics
  • Transportation

Additional details

Minimum education
Bachelor’s degree
Minimum experience
5+ years
Posting language
English
Working hours
40 hours per week