Back to job search
P
PolarGridVerified Job Source

Inference Optimization Engineer

Optimize real-time AI inference pipelines to improve latency, throughput, and cost per token. Build automated benchmarking tools and implement quantization strategies to move models from baseline to production configurations.

  • Remote
  • Canada
  • Posted Jul 28, 2026
  • 1 position

Job summary

About the Role PolarGrid is building the infrastructure layer for real-time AI inference. We're looking for an Inference Optimization Engineer to squeeze every bit of performance out of our stack. You'll work directly on the systems that serve inference to our customers, making them faster, cheaper, and more efficient. This is a deep technical, systems-focused role. You'll own latency, throughput, and cost per token as real metrics you're responsible for improving. You'll build the repeatable benchmarking and optimization process that takes new models and hardware from an initial baseline to a validated production configuration. What You'll Do Profile and optimize inference pipelines end to end using representative customer workloads, from request handling and scheduling through distributed GPU execution Tune serving frameworks such as vLLM, TensorRT-LLM, and SGLang for specific latency, throughput, and cost targets Build automated benchmarking and performance regression tooling across models, frameworks, precisions, hardware, and workload profiles Characterize customer workloads and translate TTFT, ITL, concurrency, and context-length requirements into deployment configurations Implement and evaluate quantization strategies across model families, measuring both performance gains and model-quality regressions Work with hardware teams to match model configurations and parallelism strategies to GPU topology, NVLink, and interconnect bandwidth Benchmark new hardware such as RTX Pro 6000s and B300s, identifying the best engine, precision, parallelism, and deployment configuration for each workload Bring new model architectures into production, including checkpoint conversion, framework support, distributed configuration, and correctness validation Contribute to continuous batching, speculative decoding, KV-cache optimization, prefill/decode disaggregation, and request-scheduling work Read, debug, and modify inference framework internals when configuration-level tuning is not enough Work with the platform team to canary performance improvements, measure them under production traffic, and turn successful configurations into repeatable deployment recipes What We're Looking For Strong GPU systems fundamentals, with the ability to work across Python, C++, CUDA, or Triton when optimization requires going below framework configuration Hands-on experience with at least one major inference serving framework such as vLLM, TGI, TensorRT-LLM, or SGLang ⁠Deep understanding of transformer architecture and where inference bottlenecks actually live Ability to read, debug, and modify inference framework internals rather than treating them as black boxes Experience building benchmarking, load-generation, or performance-regression infrastructure Comfortable profiling with Nsight Systems, Nsight Compute, PyTorch Profiler, or similar tools Experience with quantization and precision tradeoffs in production, including validating numerical correctness and model quality Experience optimizing multi-GPU or multi-node inference across high-speed interconnects Understanding of distributed inference, NCCL, GPU topology, and communication bottlenecks ⁠You care about numbers: TTFT, ITL, P95/P99 latency, throughput, GPU utilization, and tokens/sec/dollar Bonus Points Experience writing custom CUDA or Triton kernels Familiarity with speculative decoding, MoE routing optimizations, or prefill/decode disaggregation Experience with inference request routing, scheduling, or admission control ⁠Experience upstreaming performance improvements to vLLM, SGLang, TensorRT-LLM, or related projects ⁠Open-source contributions to inference or ML systems projects Why PolarGrid You'll work on real hardware at scale, not toy benchmarks. The performance improvements you ship go directly to customers and directly affect our unit economics. Small team, real ownership.

What you’ll do

Optimize real-time AI inference pipelines to improve latency, throughput, and cost per token. Build automated benchmarking tools and implement quantization strategies to move models from baseline to production configurations.

Requirements

Requires strong GPU systems fundamentals and hands-on experience with inference serving frameworks like vLLM or TensorRT-LLM. Candidates must be proficient in profiling tools and have a deep understanding of transformer architectures and distributed inference.

Other relevant skills

Identified from the job description. Confirm important requirements above.

  • GPU Optimization
  • vLLM
  • TensorRT-LLM
  • SGLang
  • CUDA
  • Triton
  • C++
  • Python
  • Quantization
  • Distributed Inference
  • PyTorch Profiler
  • Nsight Systems
  • Transformer Architecture
  • Benchmarking
  • NCCL
  • GPU Topology

Job areas

  • Software
  • Technology
  • Engineering
  • Data & Analytics
  • Science & Research

Additional details

Minimum experience
5+ years
Posting language
English
Working hours
40 hours per week
Location requirements
Country, Canada
Application method
Direct apply is available