Back to job search
J
Jobright.aiVerified Job Source

Site Reliability Engineer, New Grad

The role focuses on maintaining the reliability, scalability, and observability of an AI-powered platform through automation and monitoring. Key tasks include defining SLIs/SLOs, investigating system failures, and partnering with engineers for safe production releases.

  • Remote
  • Canada
  • Posted Aug 6, 2026
  • 1 position

Job summary

Jobright is your personal AI job search agent that transforms the way you do job search from solo, time-consuming efforts to a fast, expert-guided journey, simplifying every job search step and accelerating your route to the best job outcomes. The New Grad Site Reliability Engineer will help keep the AI-powered platform reliable, scalable, and observable by automating operations and improving production reliability. Why Join Us • Build real, production AI agents used by real users • High ownership and impact • Work at the intersection of AI, agents, and product • Shape how people experience AI-driven job search Responsibilities • Help define and monitor service-level indicators and objectives • Build dashboards, alerts, and reliability reporting • Automate operational tasks and reduce manual intervention • Investigate incidents, performance issues, and system failures • Improve service scalability, resilience, and recovery procedures • Partner with engineers on production readiness and safe releases • Document runbooks and contribute to incident reviews Qualification Required • Recent graduate in Computer Science, Engineering, Information Systems, or a related field • Programming or scripting experience with Python, Go, Bash, or a similar language • Familiarity with Linux, networking, databases, and distributed-system fundamentals • Strong analytical and debugging skills • Ability to communicate clearly in a remote environment Preferred • Experience with cloud platforms, containers, or Kubernetes • Familiarity with metrics, logs, traces, and observability tools • Understanding of SRE concepts such as SLIs, SLOs, and error budgets • Previous internship or systems-focused project experience

What you’ll do

The role focuses on maintaining the reliability, scalability, and observability of an AI-powered platform through automation and monitoring. Key tasks include defining SLIs/SLOs, investigating system failures, and partnering with engineers for safe production releases.

Requirements

Candidates must be recent graduates in Computer Science or a related field with programming experience in Python, Go, or Bash. Familiarity with Linux, networking, and distributed systems is required, while experience with Kubernetes and cloud platforms is preferred.

Listed skills

  • KubernetesPreferred
  • GoPreferred
  • LinuxPreferred
  • PythonPreferred

Other relevant skills

Identified from the job description. Confirm important requirements above.

  • Python
  • Go
  • Bash
  • Linux
  • Networking
  • Databases
  • Distributed Systems
  • Kubernetes
  • Cloud Platforms
  • Observability Tools
  • SRE Concepts
  • Debugging

Job areas

  • Technology
  • Software
  • Engineering
  • Data & Analytics

Additional details

Minimum education
Bachelor’s degree
Minimum experience
0+ years
Posting language
English
Working hours
40 hours per week
Seniority
Entry level