Find your next opportunity
Showing jobs within 50 km of Hamilton, ON. You can adjust the distance at any time.
Rose International
$70–$76 / hour
Job summary
Lead L3 triage and resolution of complex production incidents across the Enterprise Risk Management platform while owning the problem record lifecycle. Coordinate cross-pillar DevOps initiatives and ensure the stability, resilience, and data integrity of critical data pipelines and reporting layers.
Job details
Date Posted: 07/10/2026 Hiring Organization: Rose International Position Number: 504024 Industry: Financial Services Job Title: L3 Production Support Engineer Job Location: Mississauga, ON, Canada, L5B 1A7 Work Model: Hybrid Work Model Details: Hybrid -3 days onsite and 2 days remote Shift: Standard Work Hours Employment Type: Temp to Hire FT/PT: Full-Time Estimated Duration (In months): 6 Min Hourly Rate($): 70.00 Max Hourly Rate($): 76.00 Must Have Skills/Attributes: API, Application Support, Banking/Financial, Data Architecture, DevOps, Incident Management, Kafka, Oracle, Production Support, Stakeholders Experience Desired: Production engineering, L3 support, or platform DevOps experience in enterprise environments (8+ yrs); Experience supporting enterprise data platforms, incident, and production stability initiatives (5+ yrs); Experience leading DevOps, release management, and cross-functional technical operations teams (8+ yrs) Required Minimum Education: Bachelor’s Degree **C2C is not available** Job Description Required Education Bachelor's degree or equivalent experience in Computer Science, Engineering, Information Technology, Data Engineering, or a related field. Required Qualifications/Skills/Experience 8+ years of technology experience, with a strong background in production engineering, L3 support, data platform operations, or platform DevOps in a large-scale enterprise environment. Proven experience in a senior technical operations or engineering lead role with accountability for production stability, data pipeline health, and team coordination. Demonstrated ability to triage and resolve complex, multi-system production issues across distributed microservices and data pipeline architectures. Strong hands-on experience with data platform operations — including event-driven pipelines, data reconciliation processes, and multi-layer reporting architectures. Working knowledge of data modeling concepts — entity relationships, canonical data models, and schema evolution. Experience with data lineage analysis and tooling — tracing and documenting data flows from source systems through transformation layers to downstream consumers and identifying break points across multi-system architectures. Strong experience with ITSM processes (ServiceNow — incident, problem, change, MTP/PRJ modules) in a formal IT governance environment. Experience coordinating production releases — runbook execution, health validation, and stakeholder communication. Excellent communication, escalation management, and stakeholder engagement skills — comfortable representing technical and data quality status to senior leadership. Kafka / JMS — producer/consumer health monitoring, topic-level triage, schema validation, consumer lag analysis. Oracle — read-level query capability; working knowledge of materialized views, batch jobs, schema structures, and data reconciliation views. Data contract principles — producer/consumer responsibilities, schema validation, null-safety, field-level contract analysis. Tableau / Superset — operational dashboard monitoring and data layer reconciliation. Elastic Search — basic operational awareness for search/index layer triage. MongoDB / Couchbase — operational awareness for NoSQL data stores in use across Pillars. Data Lake concepts — retention policies, data classification, downstream consumer patterns. OpenShift / Kubernetes — pod management, health checks, container operations. Harness — deployment pipeline management and release validation. Lightspeed (LSE / Classic) — CI/CD platform management. GitHub / Bitbucket — repository governance, branch management, pipeline configuration. AppDynamics, Splunk, ELK / Kibana — application performance monitoring, log analysis, and alerting. Java / Spring Boot — sufficient depth to triage application-layer issues, interpret stack traces, and understand data contract failures. REST APIs — API failure interpretation, connectivity validation, integration troubleshooting. Angular / React — basic familiarity for front-end issue triage. SonarQube, Snyk, Checkmarx — compliance gate interpretation. CyberArk, CISAR — FID and privileged access management. EEMS / EERS — entitlement management and access review processes. CVM / CAMP — vulnerability management and Pillar-level remediation coordination. Hashicorp Vault — secrets management operations. SSL / TLS certificate lifecycle management. Preferred Qualifications/Skills/Experience Experience in financial services or regulated technology environments strongly preferred. Experience with ERDL or equivalent enterprise risk data layer platforms — aggregating data from multiple upstream systems into a central risk reporting layer. Familiarity with Kafka schema governance and event-driven integration patterns between enterprise risk systems (limits, thresholds, KRIs, model risk). Experience supporting or onboarding agentic AI or GenAI-integrated applications (e.g., Generative AI overlays, LLM-backed workflows, MCP-based architectures). Experience implementing automated data quality monitoring and self-healing alerting frameworks. Knowledge of FAST automation framework, contract testing, or behavior-driven development (Gherkin/Cucumber/Selenium). Familiarity with cloud modernization (Cloud @ Client / Type A migration) and container-native platform evolution. Exposure to enterprise risk management concepts — 1LOD/2LOD governance, stress testing (CCAR/QMMF), model risk management, or limits and thresholds frameworks. Experience with Redis infrastructure for caching and entitlement stability. Overview This is a senior technology operations leadership role responsible for the stability, resilience, data integrity, and continuous improvement of the Enterprise Risk Management (ERM) platform — a high-criticality, large-scale application comprising 17+ sub-pillars and a growing portfolio of agentic AI modules. The role sits at the intersection of production support, platform engineering, DevOps, and data platform operations, requiring deep technical capability across application support, event-driven data pipelines, and enterprise data architecture combined with strong operational discipline and cross-team coordination. The successful candidate will lead an L3 engineering function, drive DevOps maturity, own data platform operations (including ERDL and the ERM Data Lake), and serve as the technical bridge between development teams, data engineering, infrastructure, and business stakeholders. This is not a passive support role — it is an active platform ownership position with accountability for production health, data quality, release governance, incident resolution, and engineering excellence. Job Duties Lead L3 triage and resolution of complex production incidents across a 17+ pillar enterprise risk platform, including data ingestion failures, workflow disruptions, Kafka messaging issues, and infrastructure events. Own the Problem Record (PRB) lifecycle — from triage and root cause analysis to fix coordination and post-incident documentation — in alignment with ServiceNow ITSM processes. Drive the SWAT process: daily review of open problem tickets with escalation potential, ensuring senior stakeholder awareness and timely resolution. Serve as the primary escalation point for L3 engineers, coordinating with development, middleware, DBA, and Tenant Ops teams as needed. Lead Major Incident Management (MIM) for high-impact production events including data pipeline failures, SLA misses, and reconciliation discrepancies. Own L3 production support for the Enterprise Risk Data Layer (ERDL) — the central data platform that aggregates risk data from 17+ upstream Pillar systems and Federated Limits Units (FLUs). Triage and resolve data pipeline failures across Kafka, Oracle, and the ERM Data Lake — including ingestion errors, materialized view refresh failures, batch job timeouts, and data reconciliation discrepancies. Perform and govern daily data quality checks: KRI count reconciliation (ERDL vs. KRI API vs. KRI/RMD Dashboard), RA count reconciliation, PRIVATE_IND NULL checks, full and incremental refresh status monitoring, and Kafka ingestion lag monitoring. Maintain deep working knowledge of the ERDL reporting layer hierarchy and its authoritative use cases: RMD (Risk Management Dashboard) for business user consumption and management reporting; ERDL Recon View for reconciliation and data quality checks; ERDL Tableau Extract for historical as-of views for reconciliation purposes. Identify, classify, and escalate SLA Miss PRBs — distinguishing ERM application defects from upstream FLU compliance failures, and maintaining a running log of SLA miss frequency by FLU for governance escalation. Support and govern the ERM Data Lake — understanding data flows, retention policies, and downstream consumer dependencies. Maintain and apply working knowledge of the ERM canonical data model — understanding key entities (overlays, limits, thresholds, KRIs, risk assessments), their relationships, and how they flow through the platform. Understand and document end-to-end data lineage for critical data flows: from upstream FLU systems through Kafka, into ERDL, through to downstream consumers (RMD, Tableau, Superset, Recon). Apply data contract principles to triage integration failures — identifying where producer/consumer misalignments (missing fields, incorrect data types, null values, schema mismatches) are causing production defects. Contribute to data architecture governance: support the validation of AI-generated data contract analyses (e.g., Devin-produced analyses), participate in Data Contract Review meetings, and help translate findings into Jira backlog items. Classify production problems accurately as either Implementation Bugs (coding/configuration errors) or Policy/Process Misalignments (failures to correctly execute business rules), articulating the distinction clearly in PRB documentation. Support the ERDL Refactoring initiative and related data architecture modernization efforts, including Redis infrastructure upgrades and Prod Parallel Deployment stabilization. Perform and coordinate DevOps functions across lower environments (DEV, SIT, UAT), including environment management, deployment sequencing, and release validation. Lead or coordinate weekly production releases — owning the release bridge, executing health validation in Harness and OpenShift, and ensuring complete post-deployment documentation. Drive CI/CD pipeline health across Lightspeed and GitHub — ensuring build integrity, scan compliance (Snyk, SonarQube, Checkmarx), and deployment readiness across all Pillars. Manage Continuous Vulnerability Management (CVM) across the platform, coordinating remediation plans with development teams by Pillar. Lead cross-pillar DevOps initiatives: Angular/React upgrades, Python version decommissions, Tomcat upgrades, Hashicorp Vault onboarding, SSL certificate lifecycle management, and GitHub migration. Own and evolve the daily operational monitoring framework — Kafka ingestion health, data reconciliation, refresh status, and data quality checks. Drive automation of manual monitoring activities, including automated ServiceNow incident creation for detected anomalies (e.g., via Superset or alerting). Leverage Kafka dashboards (Tableau/Superset) and OpenShift tooling for real-time platform health visibility. Identify and close monitoring gaps — including proactive FID/AD group membership monitoring to prevent silent infrastructure failures. Enforce L3 operational standards: PRB description quality, PTASK lifecycle compliance, and Manual Touch Point (MTP) process adherence. Champion Developer Manifesto compliance: README standards, GitCode ownership, branch hygiene, stale repository cleanup, and CI/CD health metrics. Ensure compliance with technology risk, security, IS assessment, and regulatory standards (GIAM, EERS, CVM, CAMP, DPS data protection standards). Apply the AI-Assisted Analysis and Governance framework to accelerate root cause identification and translate findings into auditable Jira backlogs. Serve as the operational owner for new application onboarding (e.g., OMAI Overlay, Tapas, Shock Generation AI) — defining support models, escalation matrices, and runbooks. Represent BAU production health and data platform status at weekly governance calls, cross-pillar bi-weekly calls, and SWAT touchpoints. Partner with Architecture, Product Owners, Development Leads, Data Engineers, and Infrastructure teams to coordinate cross-pillar initiatives. Mentor L3 engineers; foster a culture of operational excellence, data quality ownership, and structured knowledge sharing. #CT1 **Only those lawfully authorized to work in the designated country associated with the position will be considered.** **Please note that all Position start dates and duration are estimates and may be reduced or lengthened based upon a client’s business needs and requirements.** Benefits For information and details on employment benefits offered with this position, please visit here. Should you have any questions/concerns, please contact our HR Department via our secure website. California Pay Equity For information and details on pay equity laws in California, please visit the State of California Department of Industrial Relations' website here. Rose International is an Equal Opportunity Employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, age, sex, sexual orientation, gender (expression or identity), national origin, arrest and conviction records, disability, veteran status or any other characteristic protected by law. Positions located in San Francisco and Los Angeles, California will be administered in accordance with their respective Fair Chance Ordinances. If you need assistance in completing this application, or during any phase of the application, interview, hiring, or employment process, whether due to a disability or otherwise, please contact our HR Department. Rose International has an official agreement (ID #132522), effective June 30, 2008, with the U.S. Department of Homeland Security, U.S. Citizenship and Immigration Services, Employment Verification Program (E-Verify). (Posting required by OCGA 13/10-91.).
What you’ll do
Lead L3 triage and resolution of complex production incidents across the Enterprise Risk Management platform while owning the problem record lifecycle. Coordinate cross-pillar DevOps initiatives and ensure the stability, resilience, and data integrity of critical data pipelines and reporting layers.
Requirements
Requires 8+ years of technology experience in production engineering, L3 support, or platform DevOps within large-scale enterprise environments. Candidates must hold a Bachelor's degree and possess strong hands-on experience with data platform operations, incident management, and distributed microservices architectures.
Listed skills
- Java · Preferred
- Spring Boot · Preferred
- Kubernetes · Preferred
- SQL · Preferred
Other relevant skills
Identified from the job description. Confirm important requirements above.
- API
- Application Support
- Banking
- Data Architecture
- DevOps
- Incident Management
- Kafka
- Oracle
- Production Support
- ServiceNow
- Data Engineering
- Java
- Spring Boot
- Kubernetes
- OpenShift
- SQL
Job areas
- Technology
- Finance & Accounting
- Software
- Data & Analytics
- Customer Service & Support
