Senior Site Reliability Engineer (SRE) – CDN & Web Performance
- Toronto, ON
- On-site
- Posted Oct 9, 2026
- 1 position
Opens LinkedIn
- Employment type
- Full-time
- Experience level
- Senior · 5+ years
- Posting language
- English
- Working hours
- 40 hours per week
- Seniority
- Mid-Senior level
- Application method
- Direct apply is available
Job summary
Own the reliability, monitoring, incident response, release controls, and resilience of a public-facing web platform delivered through AEM Edge Delivery Services, with a focus on CDN and edge configuration, integrations, and customer experience. Establish SRE practices including performance objectives, observability, recovery exercises, vendor escalation processes, operational documentation, and compliance evidence.
Job details
About the Role We are seeking a Senior Site Reliability Engineer (SRE) to help build and operate a modern, public-facing enterprise web platform delivered through Adobe Experience Manager (AEM) Edge Delivery Services (EDS). This is a unique reliability engineering opportunity focused on CDN and edge reliability, web performance, observability, release engineering, incident management, resilience, and third-party integrations rather than traditional server administration. You will help establish the reliability practice from the ground up and work closely with platform vendors, cybersecurity, network/CDN, privacy, production support, engineering, and enterprise technology teams in a highly regulated environment. In this role, web performance is treated as a core product requirement. You will define measurable performance objectives, monitor real-user experience, manage performance budgets, and drive improvements across the entire web delivery stack. Key Responsibilities : Application Support & Incident Management Own the operational health and monitoring of the public-facing web platform across CDN, edge configuration, DNS, TLS certificates, caching, cache invalidation, WAF, bot management, and third-party integrations. Monitor critical customer-facing journeys and identify reliability or performance issues before they significantly impact users. Build and maintain external synthetic monitoring across key templates, languages, regions, and customer journeys. Develop automated smoke tests and dependency checks to validate critical interfaces following production changes. Participate in an on-call rotation and serve as a senior technical escalation point for customer-facing incidents. Lead incident response, troubleshooting, stakeholder communication, root-cause analysis, and post-incident reviews. Establish and maintain vendor escalation procedures, including severity mapping, escalation contacts, evidence collection, SLA/OLA tracking, and incident follow-up. Create diagnostics and operational runbooks for edge-specific incidents such as content publishing failures, cache invalidation issues, DNS/TLS problems, WAF issues, and content-source outages. Change & Release Reliability Design and operate change-management processes for a Git-based production environment. Establish reliable controls around Git-based deployments while maintaining enterprise change-management and audit requirements. Own CI/CD quality and security controls, including branch protection, required checks, linting, automated testing, performance validation, and secret scanning. Establish repeatable rollback, revert, republish, and cache-purge procedures. Rehearse recovery procedures and measure recovery time to ensure rollback processes are operationally proven. Support high-volume content publishing with appropriate approval, attribution, traceability, and audit evidence. Represent the platform through change governance and release-management processes. Maintain release documentation, freeze calendars, deployment procedures, and operational readiness requirements. Business Continuity & Resilience Define and test recovery strategies for the content source, Git repository, CDN configuration, and edge platform configuration. Manage CDN and edge configuration using Infrastructure as Code (IaC) / Configuration as Code principles wherever practical. Define Recovery Time Objectives (RTO) and Recovery Point Objectives (RPO) in partnership with business and technology stakeholders. Plan and execute disaster recovery, resilience, and operational continuity exercises. Test realistic failure scenarios, including: Certificate expiration DNS issues Cache invalidation failures WAF misconfiguration Content-source unavailability Repository issues or compromise CDN/edge configuration failures Maintain operational resilience documentation and evidence required for a regulated enterprise technology environment. Identify platform dependencies and maintain appropriate recovery and exit/portability strategies for critical third-party services. Reliability & Web Performance Engineering Define and maintain Service Level Indicators (SLIs), Service Level Objectives (SLOs), and error budgets for availability, reliability, and web performance. Establish Core Web Vitals targets and performance thresholds across key page templates and customer journeys. Build an observability practice using Real User Monitoring (RUM), CDN/edge logs, enterprise SIEM, metrics, and external synthetic monitoring. Analyze web performance regressions using browser/network telemetry and identify the specific source of degradation. Review browser waterfalls, rendering behavior, network requests, JavaScript execution, CSS, third-party scripts, APIs, and other dependencies. Establish governance around third-party scripts and tags, including performance measurement, approval processes, and performance budgets. Monitor CDN usage, asset storage, media delivery, egress, and other platform-related capacity and cost considerations. Identify trends and recurring reliability issues and drive preventative engineering improvements. Produce reliability, availability, and web-performance reporting for technology, risk, and business leadership. Compliance & Operational Controls Establish operational controls and evidence for a vendor-operated / SaaS-based technology platform. Support SIEM log ingestion, retention, access reviews, repository controls, configuration management, and audit evidence. Partner with cybersecurity, privacy, compliance, risk, and third-party risk teams. Maintain accurate configuration management records, support models, ownership, assignment groups, and operational documentation. Ensure reliability and operational practices align with enterprise security, regulatory, and third-party risk requirements. Support technology control assessments and provide operational evidence for audits and regulatory reviews. What You'll Build : This is a build-and-run SRE role. During the initial phase, you will help establish and mature the reliability engineering practice, including: Real User Monitoring (RUM) and Core Web Vitals dashboards CDN and edge log ingestion and alerting External synthetic monitoring CDN and edge configuration as code Reliability and operational runbooks Incident response and escalation playbooks Vendor escalation and operational support processes Git-based change and release controls Automated smoke and dependency testing Disaster recovery and resilience procedures Rehearsed rollback and recovery processes SLOs, SLIs, error budgets, and reliability reporting Operational readiness and compliance evidence Required Qualifications : Strong hands-on experience operating high-traffic, public-facing websites behind an enterprise CDN. Deep expertise in CDN configuration and edge delivery, including: Origin behavior Caching strategies Cache invalidation Edge logic TLS/SSL DNS Routing and traffic behavior Strong experience with Akamai and/or Cloudflare is highly desirable. Practical experience with Web Application Firewalls (WAF), including rule management, tuning, false-positive analysis, and production enforcement. Experience with bot management and protecting web applications while allowing legitimate search-engine and business-critical crawlers. Strong observability experience using logs, metrics, Real User Monitoring, synthetic monitoring, SLIs, SLOs, and error budgets. Strong understanding of web performance and Core Web Vitals, including browser rendering, network behavior, page-load performance, and waterfall analysis. Comfortable reviewing and troubleshooting JavaScript and CSS in modern web applications. Strong experience with Git-based development and release engineering. Hands-on experience with CI/CD pipelines and production deployment controls. Experience with Infrastructure as Code / Configuration as Code. Proven experience supporting or leading customer-facing incident management and incident response. Strong understanding of root-cause analysis, post-incident reviews, operational readiness, and continuous reliability improvement. Experience working within enterprise change management, audit, access-control, compliance, and third-party risk processes. Ability to work effectively across engineering, security, infrastructure, vendors, production support, and business stakeholders. Preferred Qualifications : Experience operating vendor-managed or SaaS platforms, where reliability requires monitoring, escalation, governance, and vendor accountability. Experience with Adobe Experience Manager (AEM), particularly AEM Edge Delivery Services or Assets as a Cloud Service. Experience in financial services, banking, healthcare, or another highly regulated industry. Experience supporting multilingual or bilingual public-facing websites. Knowledge of web accessibility and SEO from a reliability, performance, compliance, and discoverability perspective. Automation experience using Python, Java, or JavaScript. Experience building automated operational tooling, diagnostics, synthetic tests, or runbooks. Experience with enterprise SIEM and log-management platforms. Experience defining and reporting reliability and performance metrics to senior technology stakeholders. What This Role Is Not : This role is specifically focused on edge reliability, CDN, web performance, observability, release engineering, incident management, and modern platform reliability. It does not primarily involve: Operating-system patching Kubernetes or cluster administration Container orchestration Traditional application-server administration Traditional database administration Server-based infrastructure operations Traditional compute capacity planning The underlying delivery infrastructure is largely operated by the platform provider. Your focus will instead be on the reliability surface owned by the enterprise, including CDN/edge configuration, content delivery, integrations, release processes, observability, performance, resilience, and operational controls. Why Join This Team? This is an opportunity to help define reliability engineering for a modern, serverless/edge-based enterprise web platform rather than simply operating an existing infrastructure environment. You will have the opportunity to: Build a reliability practice from the ground up Work with modern CDN and edge technologies Own enterprise-grade web performance and Core Web Vitals Establish SRE principles including SLOs, SLIs, and error budgets Build observability around RUM, synthetics, and edge telemetry Lead customer-facing incident response Design resilience and disaster-recovery practices Work directly with technology vendors and enterprise stakeholders Influence how reliability, performance, and operational excellence are implemented across a high-visibility public platform If you are an experienced SRE / Reliability Engineer with deep CDN, edge, web-performance, observability, and incident-management experience, we would like to hear from you.
What you’ll do
Own the reliability, monitoring, incident response, release controls, and resilience of a public-facing web platform delivered through AEM Edge Delivery Services, with a focus on CDN and edge configuration, integrations, and customer experience. Establish SRE practices including performance objectives, observability, recovery exercises, vendor escalation processes, operational documentation, and compliance evidence.
Requirements
Requires substantial hands-on experience operating high-traffic websites behind enterprise CDNs, with expertise in edge delivery, caching, DNS, TLS, WAF, web performance, observability, and customer-facing incident response. Candidates should also have Git-based release and CI/CD experience, Infrastructure as Code knowledge, and experience working with enterprise change management, audit, security, and third-party risk controls.
Listed skills
- Incident Management · Preferred
- CI/CD · Preferred
Other relevant skills
Identified from the job description. Confirm important requirements above.
- CDN And Edge Reliability
- Web Performance Engineering
- Incident Management
- Observability
- Core Web Vitals
- Real User Monitoring
- Synthetic Monitoring
- SLOs, SLIs, And Error Budgets
- Akamai And Cloudflare
- Web Application Firewall Management
- Git-Based Release Engineering
- CI/CD
- Infrastructure As Code
- Disaster Recovery And Resilience
- JavaScript And CSS Troubleshooting
- Enterprise Compliance And Audit Controls
Job areas
- Technology
- Software
- Engineering
- Security & Safety
More jobs from N2P Systems
Power Platform & GenAI Developer
- On-site
- Montréal, QC
- Posted Oct 8, 2026
Pega Technical Architect
- On-site
- Mississauga, ON
- Posted Oct 8, 2026
Senior QA Automation Engineer – Contact Center
- On-site
- Toronto, ON
- Posted Oct 1, 2026
