Back to job search
S
SumerSportsVerified Job Source

Data Engineer

You will design, build, and maintain robust data pipelines for ingestion, transformation, and orchestration to support AI and analytics applications. You will also collaborate with ML/AI teams to develop retrieval pipelines and ensure data reliability, scalability, and compliance.

  • Remote
  • Canada, United States
  • Posted Aug 13, 2026
  • 1 position

More jobs you can apply to directly

Similar opportunities posted by employers hiring on Jobs.ca, with no external application form.

Job summary

Position Summary Our data engineering team is the foundation of this system, ensuring data is accurate, fast, and always available for our models and AI applications. As a Data Engineer, you’ll design, build, and maintain the data pipelines that power our deep learning, video and LLM systems. You’ll work across ingestion, transformation, and orchestration layers — from real-time feeds to analytics-ready datasets. Your mission is to make data reliable, discoverable, and scalable for use by model training, analytics, and AI-driven products across multiple sports. You’ll collaborate closely with our MLOps, and Sports Data teams to ensure seamless integration between data and AI. Responsibilities Build and operate robust data pipelines for ingestion, cleaning, and transformation using Databricks, Airflow, or Kubernetes. Develop efficient ETL/ELT workflows in Python and SQL to support both batch and streaming workloads. Partner with ML/AI teams to make datasets and tools discoverable and safe for autonomous agents, including evaluation and guardrails for AI-generated queries. Develop retrieval pipelines (RAG, vector search) over structured stats and unstructured sources (scouting notes, video metadata) to power AI applications. Model and maintain structured data assets (Delta, Parquet, Iceberg) for reliability, versioning, and lineage tracking. Implement orchestration and monitoring: schedule jobs, track dependencies, and automate recovery from failures. Ensure data quality and compliance through validation frameworks, schema enforcement, and audit logging. Contribute to data platform evolution: evaluate tools, standardize best practices, and improve developer experience. Support performance and cost optimization across compute, storage, and orchestration systems. Qualifications 3–8 years of experience as a Data Engineer or ETL Developer in a production environment. Proficiency in Python and SQL; strong familiarity with Databricks, Spark, or equivalent big-data frameworks. Experience with workflow orchestration tools such as Airflow, Dagster, Luigi or Prefect. Deep understanding of data modeling, data warehousing, and distributed data processing. Knowledge of modern data lakehouse architectures. Familiarity with CI/CD, GitHub Actions, Infrastructure as Code, and data pipeline testing frameworks. Comfort working in a cross-functional environment with ML, product, and analytics teams. Exposure to LLM-powered data tools: text-to-SQL, RAG, agent/tool interfaces (e.g. MCP), or natural-language analytics. Previous work with cloud infrastructure (AWS, GCP, or Azure) and container orchestration (Docker, Kubernetes). Preferred Previous experience with sports, telemetry, or sensor data pipelines. Familiarity with streaming frameworks and event driven data processing (Kafka, Spark Structured Streaming, Flink). General knowledge of American football, the NFL, and college football. Background in data governance, lineage, and observability tools (Monte Carlo, Great Expectations, Unity Catalog, OpenLineage). Experience designing semantic layers or metric definitions consumed by AI and BI tools. Exposure to best practices in machine-learning model management and MLOps. Benefits Competitive Salary and Bonus Plan Comprehensive health insurance plan Retirement savings plan (401k) with company match Remote working environment A flexible, unlimited time off policy Generous paid holiday schedule - 13 in total including Monday after the Super Bowl SumerSports is committed to fair and equitable compensation practices. Actual compensation packages are based on several factors that are unique to each candidate, including but not limited to skill set, depth of experience, certifications, and specific work location. This may be different in other locations due to differences in the cost of labor. The total compensation package for this position may also include annual performance bonus, benefits and/or other applicable incentive compensation plans.

What you’ll do

You will design, build, and maintain robust data pipelines for ingestion, transformation, and orchestration to support AI and analytics applications. You will also collaborate with ML/AI teams to develop retrieval pipelines and ensure data reliability, scalability, and compliance.

Requirements

Candidates must have 3–8 years of experience in data engineering with proficiency in Python, SQL, and big-data frameworks like Spark or Databricks. Strong familiarity with workflow orchestration tools, cloud infrastructure, and modern data lakehouse architectures is required.

Benefits

• Competitive Salary • Bonus Plan • Comprehensive Health Insurance • Retirement Savings Plan (401k) • Company Match • Remote Working Environment • Unlimited Time Off Policy • Paid Holiday Schedule

Listed skills

  • PythonPreferred
  • SQLPreferred
  • KubernetesPreferred
  • CI/CDPreferred

Other relevant skills

Identified from the job description. Confirm important requirements above.

  • Python
  • SQL
  • Databricks
  • Spark
  • Airflow
  • Kubernetes
  • ETL/ELT
  • Data Modeling
  • Data Warehousing
  • RAG
  • Vector Search
  • CI/CD
  • GitHub Actions
  • Infrastructure as Code
  • Cloud Infrastructure
  • Streaming Frameworks
  • Sensor Data
  • Pipelines
  • MLOps (Machine Learning Operations)
  • Apache Parquet
  • Observability
  • Workflow Management
  • Apache Airflow
  • Schema Markup
  • Infrastructure as Code (IaC)
  • Data Lakehouse
  • Artificial Intelligence
  • Applications Of Artificial Intelligence
  • Amazon Web Services
  • Auditing
  • Microsoft Azure
  • Business Intelligence
  • Big Data
  • Management
  • Data Processing
  • Data Engineering
  • Data Governance
  • Extract Transform Load (ETL)
  • Data Quality
  • Development Environment
  • Distributed Data Store
  • Event-Driven Programming
  • Github
  • Scalability
  • Python (Programming Language)
  • Machine Learning
  • Microsoft Certified Professional
  • Metadata
  • Telemetry
  • American Football

Job areas

  • Data & Analytics
  • Software
  • Technology
  • Engineering
  • Sports & Recreation
  • Data Engineer
  • Software Developers
  • Database Administrators

Additional details

Minimum experience
5+ years
Posting language
English
Working hours
40 hours per week