Back to job search
SA
SoundHound AIVerified Job Source

Staff Site Reliability Engineer

You will design, build, and maintain highly available and scalable infrastructure on Google Cloud Platform while automating CI/CD pipelines. Additionally, you will lead incident response, optimize system performance, and drive department-wide compliance initiatives.

  • Remote
  • Canada
  • Posted Aug 7, 2026
  • 1 position

Job summary

THE OPPORTUNITY We’re looking for a Staff Software Engineer (SRE) to join our Retail and Restaurants AI team. You will be responsible for the reliability, scalability, and performance of our infrastructure, with a deep focus on Google Cloud Platform (GCP). You will architect and maintain high-availability systems, automate operational tasks, and ensure our services can handle the demands of millions of voice AI interactions. WHAT YOU'LL DO * Design, build, and maintain highly available and scalable infrastructure on Google Cloud Platform. * Architect and automate CI/CD pipelines to ensure rapid, reliable deployments. * Implement robust monitoring, alerting, and observability strategies to proactively identify and resolve system issues. * Partner with engineering teams to optimize performance, cost, and reliability of backend services. * Drive incident response, post-mortem analysis, and long-term remediation efforts. * Identify and eliminate sources of toil, promoting operational maturity and self-service capabilities. * Collaborate with cross-functional teams to ensure alignment on infrastructure roadmaps and security standards. * Lead department wide compliance (PCI, SOC) initiatives. WHAT YOU'LL BRING * 12+ years of software engineering experience, with significant experience in Site Reliability Engineering or DevOps roles. * Expert-level experience with Google Cloud Platform (GCP) services (e.g., GKE, Compute Engine, Cloud Run, Pub/Sub). * Proficient in Infrastructure as Code (IaC) tools like Terraform or Pulumi. * Deep experience with Kubernetes, container orchestration, and service mesh architectures. * Strong background in monitoring and observability tools (e.g., Datadog, Prometheus, Grafana, Cloud Monitoring). * Experience designing and managing high-throughput, distributed systems. * Strong problem-solving skills and a growth mindset—comfortable with ambiguity and making high-stakes technical trade-offs. * Excellent communication skills and a demonstrated ability to mentor engineers. PREFERRED QUALIFICATIONS * Experience working in a high-velocity, customer-focused environment. * Familiarity with functional programming paradigms (e.g., Clojure/ClojureScript). * Prior experience in the restaurant technology, hospitality, or AI-driven SaaS space. * Experience implementing security and compliance best practices in the cloud. WORKPLACE & COMPENSATION This role is available throughout Canada. Compensation includes salary, equity, comprehensive healthcare, paid time off, and other benefits. Our recruiting team will provide a specific salary range based on location and years of experience. #LI-MQ1 #LI-REMOTE

What you’ll do

You will design, build, and maintain highly available and scalable infrastructure on Google Cloud Platform while automating CI/CD pipelines. Additionally, you will lead incident response, optimize system performance, and drive department-wide compliance initiatives.

Requirements

Candidates must have 12+ years of software engineering experience with significant expertise in SRE or DevOps roles. Expert-level proficiency in Google Cloud Platform, Kubernetes, and Infrastructure as Code tools is required.

Benefits

• Salary • Equity • Comprehensive healthcare • Paid time off

Listed skills

  • KubernetesPreferred
  • CI/CDPreferred
  • TerraformPreferred

Other relevant skills

Identified from the job description. Confirm important requirements above.

  • Google Cloud Platform
  • Site Reliability Engineering
  • DevOps
  • Kubernetes
  • Infrastructure as Code
  • Terraform
  • Pulumi
  • CI/CD
  • Observability
  • Datadog
  • Prometheus
  • Grafana
  • Distributed systems
  • Mentoring
  • Compliance
  • Cloud security
  • Pipelines
  • Cloud Administration
  • Time Off Management
  • Google Kubernetes Engine (GKE)
  • Google Cloud Platform (GCP)
  • Infrastructure as Code (IaC)
  • Artificial Intelligence
  • Customer Service
  • Clojure
  • Software As A Service (SaaS)
  • Communication
  • Incident Response
  • Scalability
  • Problem Solving
  • Publish Subscribe
  • Prometheus (Software)
  • Restaurant Operation
  • Self Service Technologies
  • Software Engineering
  • Functional Programming

Job areas

  • Technology
  • Software
  • Engineering
  • Retail
  • Hospitality
  • Staff Site Reliability Engineer
  • Site Reliability Engineer
  • Software and Applications Developers and Analysts Not Elsewhere Classified
  • Computer and Information Research Scientists

Additional details

Minimum experience
10+ years
Posting language
English
Working hours
40 hours per week