Back to job search
HeyGen logo

Software Engineer, GPU Performance

  • Toronto, ON
  • On-site
  • Posted Oct 1, 2026
  • 1 position

Opens an external site

Sign in to save this job
Employment type
Full-time
Experience level
Mid-level · 2+ years
Posting language
English
Working hours
40 hours per week

Job summary

The engineer will optimize GPU performance for AI applications by identifying bottlenecks in model execution and inference infrastructure. They will develop high-performance kernels and implement optimizations to improve latency, throughput, and GPU cost.

Job details

About HeyGen At HeyGen, our mission is to make visual storytelling accessible to all. Over the last decade, visual content has become the preferred method of information creation, consumption, and retention. But the ability to create such content, in particular videos, continues to be costly and challenging to scale. Our ambition is to build technology that equips more people with the power to reach, captivate, and inspire audiences. Learn more at www.heygen.com. Visit our Mission and Culture doc here. Position Summary HeyGen is building AI applications including Avatar IV, Photo Avatar, Interactive Avatar, and Video Translation. We’re looking for a Software Engineer focused on GPU performance to make the systems behind these experiences faster and more efficient. You will work across model execution and inference infrastructure, using profiling and measurement to improve latency, throughput, and GPU cost. This role is a fit for an engineer who enjoys understanding how software uses the hardware beneath it. Key Responsibilities Use NVIDIA Nsight Systems, Nsight Compute, and PyTorch Profiler to investigate GPU utilization, kernel execution, memory bandwidth, and CPU–GPU data movement. Identify bottlenecks across model execution, preprocessing, and inference serving, then measure the impact of each optimization. Improve performance through batching, scheduling, memory management, and better GPU utilization. Develop or integrate high-performance GPU kernels when existing implementations limit performance. Build benchmarks and automated checks that catch performance regressions across representative video workloads. Collaborate with AI researchers and infrastructure engineers to bring optimizations into production. Measure the effect of changes on latency, throughput, cost, and output quality. Qualifications Experience optimizing GPU-based AI workloads or high-performance computing systems. Proficiency in Python and experience with PyTorch or a similar machine learning framework. Strong curiosity about GPU hardware, including memory bandwidth, cache behavior, tensor cores, and data movement between CPU and GPU. Experience using profiling tools to connect hardware behavior to application-level bottlenecks and validate improvements. Ability to turn performance experiments into reliable production changes and communicate tradeoffs clearly. Preferred Qualifications Experience with CUDA, Triton, or C++ GPU programming. Experience optimizing video, image, audio, diffusion, or Transformer models. Familiarity with multi-GPU inference, GPU interconnects, quantization, or large-scale model serving. Experience building performance benchmarks or regression testing infrastructure. Prior experience in a fast-paced technology environment. What HeyGen Offers Competitive salary and benefits package. Dynamic and inclusive work environment. Opportunities for professional growth and advancement. Collaborative culture that values innovation and creativity. Access to the latest technologies and tools. HeyGen is an Equal Opportunity Employer. We celebrate diversity and are committed to creating an inclusive environment for all employees. Join us at HeyGen Help us make AI video faster and more accessible. We’d love to hear from you.

What you’ll do

The engineer will optimize GPU performance for AI applications by identifying bottlenecks in model execution and inference infrastructure. They will develop high-performance kernels and implement optimizations to improve latency, throughput, and GPU cost.

Requirements

Candidates must have experience optimizing GPU-based AI workloads and proficiency in Python with machine learning frameworks. Strong knowledge of GPU hardware architecture and experience with profiling tools are essential for this role.

Benefits

  • Competitive salary
  • Professional growth opportunities
  • Inclusive work environment
  • Access to latest technologies

Listed skills

  • Python · Preferred
  • C++ · Preferred

Other relevant skills

Identified from the job description. Confirm important requirements above.

  • GPU performance optimization
  • Python
  • PyTorch
  • NVIDIA Nsight Systems
  • NVIDIA Nsight Compute
  • PyTorch Profiler
  • CUDA
  • Triton
  • C++
  • Inference infrastructure
  • Memory management
  • Performance benchmarking
  • Kernel execution
  • Quantization
  • Transformer models
  • Avatar
  • Transformer (Machine Learning Model)
  • Workplace Inclusivity
  • Curiosity
  • Visual Storytelling
  • Research
  • Artificial Intelligence
  • Applications Of Artificial Intelligence
  • Automation
  • Building Performance
  • C++ (Programming Language)
  • Nvidia CUDA
  • Creativity
  • Extract Transform Load (ETL)
  • Memory Management
  • Innovation
  • Python (Programming Language)
  • Machine Learning
  • Regression Testing
  • Software Engineering
  • Scheduling
  • PyTorch (Machine Learning Library)

Job areas

  • Software
  • Technology
  • Engineering
  • Data & Analytics
  • Software Performance Engineer
  • Software Developer / Engineer
  • Software Developers

More jobs from HeyGen

See all jobs from HeyGen