Back to job search
Kubex logo
KubexVerified Job Source

Staff Software Engineer - Kubernetes Observability

  • Canada
  • Remote
  • Posted Aug 25, 2026
  • 1 position

Opens an external site

Sign in to save this job
Employment type
Full-time
Experience level
Senior · 5+ years
Posting language
English
Working hours
40 hours per week
Seniority
Mid-Senior level
Application method
Direct apply is available

Job summary

Lead the architecture and development of Kubernetes observability capabilities to support AI-driven infrastructure optimization. Design systems for collecting telemetry and improving the performance of inference workloads across complex customer environments.

Job details

About Kubex Kubex is building the future of autonomous, AI-driven infrastructure optimization. Our platform enables intelligent, policy-driven optimization across Kubernetes, cloud, and GPU-backed environments improving performance, reducing cost, and eliminating waste for some of the world’s most sophisticated technology organizations. As AI workloads increasingly run on Kubernetes, especially for inference at scale, Kubex is expanding its GPU support to provide more advanced optimization and automation aligned with the unique challenges of GPU-accelerated infrastructure. We combine deep systems expertise, advanced analytics, and patented optimization technology to help customers run AI workloads efficiently and reliably in real-world production environments. Role Overview Kubex is seeking a Technical Lead, Kubernetes Observability to own the architecture, technical direction and development of our Kubernetes observability capabilities. This is a senior technical leadership role for someone who can define how telemetry is collected, enriched, processed, and used to support infrastructure optimization across complex customer environments. You will lead the design and evolution of systems that capture workload behavior, resource usage, performance, and operational context across Kubernetes clusters. This includes making key architectural decisions, establishing engineering patterns, prototyping new approaches, and contributing directly to the most critical and technically challenging parts of the implementation. This role combines high-level ownership with strong hands-on engineering. You will use modern AI-assisted development workflows to accelerate implementation, testing, investigation, and iteration, while remaining accountable for system design, technical decisions, code quality, and production reliability. The ideal candidate has deep experience building observability or telemetry systems for Kubernetes and is comfortable working across metrics, events, logs, traces, workload metadata, and distributed data collection. Experience with GPU or AI workloads is valuable but not required. You will help build the observability foundation that supports Kubex’s current Kubernetes optimization capabilities and its continued expansion into GPU and AI infrastructure. Key Responsibilities Lead the design of systems that collect telemetry and improve performance of inference workloads. Contribute directly to production code, remaining deeply hands-on in the design, implementation, and evolution of core platform components. Collaborate closely with other senior engineers, product managers and engineering leadership to coordinate and execute complex software development initiatives. Prototype, validate, and productionize new technical approaches related to AI workload observability and performance optimization. Identify opportunities to extend Kubex’s value beyond inference workloads, including potential future optimizations for training or hybrid workloads. Required Qualifications 7+ years of professional software engineering experience, including significant experience building production software on Kubernetes. Strong experience designing or building observability and telemetry solutions using technologies such as Prometheus, OpenTelemetry, or similar platforms. Deep understanding of Kubernetes workloads, resource management, API interactions, and the operational challenges of running software across diverse customer clusters. Strong coding skills, preferably in Go, with experience building scalable, reliable, and testable distributed systems. Demonstrated ability to own technical architecture while remaining hands-on with implementation, prototyping, debugging, and production delivery. Preferred Qualifications Experience building collectors, exporters, agents, or telemetry-processing pipelines for Kubernetes environments. Knowledge of performance analysis, resource optimization, scheduling, or infrastructure efficiency. Exposure to GPU-backed infrastructure, AI inference workloads, or GPU observability technologies. Why Join Kubex? Play a key role in shaping the future of AI infrastructure optimization. Work on technically challenging problems at the intersection of Kubernetes, GPUs, and AI workloads. Collaborate with a highly experienced, deeply technical team. Influence product direction, architecture, and external technical positioning. Flexible, remote-first culture focused on impact and innovation. Competitive compensation, equity, and benefits.

What you’ll do

Lead the architecture and development of Kubernetes observability capabilities to support AI-driven infrastructure optimization. Design systems for collecting telemetry and improving the performance of inference workloads across complex customer environments.

Requirements

Requires over 7 years of software engineering experience with a strong background in building production software on Kubernetes and telemetry solutions. Proficiency in Go and experience with distributed systems and resource management are essential.

Benefits

• Competitive Compensation • Equity • Benefits

Listed skills

  • Kubernetes · Preferred
  • Go · Preferred

Other relevant skills

Identified from the job description. Confirm important requirements above.

  • Kubernetes
  • Observability
  • Telemetry
  • Go
  • Prometheus
  • OpenTelemetry
  • Distributed Systems
  • GPU Infrastructure
  • AI Workloads
  • Resource Management
  • System Architecture
  • Performance Optimization

Job areas

  • Software
  • Technology
  • Engineering
  • Data & Analytics

More jobs you can apply to directly

Similar opportunities posted by employers hiring on Jobs.ca, with no external application form.

Browse all Easy Apply jobs