Senior Databricks Data Engineer
Design and maintain enterprise-scale batch and streaming pipelines using a Medallion Architecture on a Databricks Lakehouse platform. Optimize Spark performance and implement data governance and security via Unity Catalog.
- On-site
- Brampton, ON
- Posted Aug 13, 2026
- Apply by Sep 12, 2026
- 1 position
More jobs you can apply to directly
Similar opportunities posted by employers hiring on Jobs.ca, with no external application form.
Forgeahead Solutions Corporation
Technical Lead and Senior Software Engineer
- On-site
BC Public Schools
Manager, People, Performance and Culture
- On-site
Alcohol and Gaming Commission of Ontario (AGCO)
Information Management Lead / Responsable de la gestion de l’information
- On-site
Job summary
Position Overview We are looking for a Senior, Super Hands-On Databricks Data Engineer who lives and breathes code, query optimization, and modern data architecture. In this role, you won't just design architectures on whiteboards—you will write production PySpark/SQL, optimize Databricks clusters, build streaming and batch pipelines, and enforce data governance. You will own end-to-end pipeline execution from raw ingestion to curated Gold layer models, playing a lead role in modernizing our Lakehouse platform. Key Responsibilities 1. Hands-On Pipeline Development & Lakehouse Architecture Design, build, and maintain enterprise-scale batch and real-time streaming pipelines using PySpark, SQL, Delta Live Tables (DLT), and Auto Loader. Implement and refine Medallion Architecture (Bronze Silver Gold) to support downstream BI, reporting, and Machine Learning workloads. Enforce schema evolution, ACID transactions, and data compaction using Delta Lake core constructs. 2. Performance Tuning & Optimization (Deep Tech) Diagnose and resolve Spark performance bottlenecks: data skew, OOM errors, excessive shufflings, and memory spills. Optimize queries using Liquid Clustering, Z-Ordering, Data Partitioning, AQE (Adaptive Query Execution), and Photon engine tuning. Benchmark and optimize Databricks compute workloads to minimize DBU (Databricks Unit) consumption and cloud costs (FinOps). 3. Governance, Security & Quality Implement end-to-end data governance, fine-grained access control (row/column-level security), and lineage tracking using Unity Catalog. Automate automated data quality validation checks and alert mechanisms across the pipeline life cycle. 4. Operations, CI/CD & DevOps Automate pipeline orchestration using Databricks Asset Bundles (DABs) or Databricks Workflows / Apache Airflow. Build CI/CD pipelines (GitHub Actions, Azure DevOps, or GitLab) for automated testing, deployment, and code promotions. Requirements Required Skills & Qualifications Must-Haves Experience: 8+ years in Data Engineering, with 4+ years of intensive, hands-on production experience on Databricks. Programming Mastery: Fluent in PySpark, Advanced SQL, and Python. Databricks Ecosystem: Deep experience with Delta Lake, Unity Catalog, Delta Live Tables (DLT), Auto Loader, and Databricks Workflows. Cloud Infrastructure: Strong hands-on experience in at least one primary cloud provider (AWS, Azure, or GCP) integration with Databricks (S3/ADLS Gen2, IAM, Key Vaults/Secret Manager). Data Modeling: Solid understanding of dimensional modeling (Kimball), One Big Table (OBT) strategies, and data vault patterns. CI/CD & Software Engineering: Proficient in Git workflows, unit testing PySpark code (pytest), and deployment automation. Preferred / Nice-to-Haves Certifications: Databricks Certified Data Engineer Professional. Streaming: Hands-on with Apache Kafka, Event Hubs, or Kinesis integration via Structured Streaming. GenAI / ML Ops: Familiarity with MLflow, Feature Store, or Vector Search within Databricks. Infrastructure as Code (IaC): Experience using Terraform to provision Databricks workspaces and storage resources. Performance Indicators (How success is measured) Pipeline Reliability: Maintaining strict SLA thresholds on critical Gold-layer models. Cost Efficiency: Measurable reduction in DBU costs through effective compute profiling and tuning. Code Quality: High test coverage and zero-downtime CI/CD deployments.
What you’ll do
Design and maintain enterprise-scale batch and streaming pipelines using a Medallion Architecture on a Databricks Lakehouse platform. Optimize Spark performance and implement data governance and security via Unity Catalog.
Requirements
Requires over 8 years of data engineering experience, with at least 4 years of intensive production experience in Databricks. Mastery of PySpark, SQL, and Python is essential, along with proficiency in CI/CD and cloud infrastructure.
Listed skills
- CI/CDPreferred
- PythonPreferred
Other relevant skills
Identified from the job description. Confirm important requirements above.
- PySpark
- Advanced SQL
- Python
- Databricks
- Delta Lake
- Unity Catalog
- Delta Live Tables
- Auto Loader
- Medallion Architecture
- Spark Performance Tuning
- CI/CD
- Data Modeling
- Cloud Infrastructure
- Apache Airflow
- GitHub Actions
- Azure DevOps
Job areas
- Data & Analytics
- Technology
- Software
- Engineering
- Transportation
Additional details
- Minimum experience
- 10+ years
- Apply by
- Sep 12, 2026
- Posting language
- English
- Working hours
- 40 hours per week
- Seniority
- Mid-Senior level
- Application method
- Direct apply is available
