Site Reliability Engineer, New Grad
The role focuses on maintaining the reliability, scalability, and observability of an AI-powered platform through automation and monitoring. Key tasks include defining SLIs/SLOs, investigating system failures, and partnering with engineers for safe production releases.
- Remote
- Canada
- Posted Aug 6, 2026
- 1 position
Job summary
Jobright is your personal AI job search agent that transforms the way you do job search from solo, time-consuming efforts to a fast, expert-guided journey, simplifying every job search step and accelerating your route to the best job outcomes. The New Grad Site Reliability Engineer will help keep the AI-powered platform reliable, scalable, and observable by automating operations and improving production reliability. Why Join Us • Build real, production AI agents used by real users • High ownership and impact • Work at the intersection of AI, agents, and product • Shape how people experience AI-driven job search Responsibilities • Help define and monitor service-level indicators and objectives • Build dashboards, alerts, and reliability reporting • Automate operational tasks and reduce manual intervention • Investigate incidents, performance issues, and system failures • Improve service scalability, resilience, and recovery procedures • Partner with engineers on production readiness and safe releases • Document runbooks and contribute to incident reviews Qualification Required • Recent graduate in Computer Science, Engineering, Information Systems, or a related field • Programming or scripting experience with Python, Go, Bash, or a similar language • Familiarity with Linux, networking, databases, and distributed-system fundamentals • Strong analytical and debugging skills • Ability to communicate clearly in a remote environment Preferred • Experience with cloud platforms, containers, or Kubernetes • Familiarity with metrics, logs, traces, and observability tools • Understanding of SRE concepts such as SLIs, SLOs, and error budgets • Previous internship or systems-focused project experience
What you’ll do
The role focuses on maintaining the reliability, scalability, and observability of an AI-powered platform through automation and monitoring. Key tasks include defining SLIs/SLOs, investigating system failures, and partnering with engineers for safe production releases.
Requirements
Candidates must be recent graduates in Computer Science or a related field with programming experience in Python, Go, or Bash. Familiarity with Linux, networking, and distributed systems is required, while experience with Kubernetes and cloud platforms is preferred.
Listed skills
- KubernetesPreferred
- GoPreferred
- LinuxPreferred
- PythonPreferred
Other relevant skills
Identified from the job description. Confirm important requirements above.
- Python
- Go
- Bash
- Linux
- Networking
- Databases
- Distributed Systems
- Kubernetes
- Cloud Platforms
- Observability Tools
- SRE Concepts
- Debugging
Job areas
- Technology
- Software
- Engineering
- Data & Analytics
Additional details
- Minimum education
- Bachelor’s degree
- Minimum experience
- 0+ years
- Posting language
- English
- Working hours
- 40 hours per week
- Seniority
- Entry level
