Opens an external site
- Employment type
- Full-time
- Experience level
- Senior · 5+ years
- Posting language
- English
- Working hours
- 40 hours per week
- Seniority
- Mid-Senior level
- Application method
- Direct apply is available
Job summary
Design, build, and maintain scalable automated data pipelines for batch and streaming ingestion using Databricks. Implement robust data models and architectures while integrating monitoring and alerting mechanisms to ensure pipeline health.
Job details
Description This job posting is for an existing, active vacancy and we are looking to hire Sr. Data Engineer in WFH (Canada), immediately who has strong experience in Databricks, Delta Live Tables, Oops, Python , PyTest and PySpark. Role Overview Mandatory Skill: Programming Languages: Proficient in Python, PyTest and PySpark with a strong understanding software engineering best practice. Oops Concept: Strong Knowledge of It. Cloud Computing: Utilize Azure cloud-based data platforms, specifically leveraging Databricks and Delta Live Tables for data engineering tasks, while effectively utilizing services related to storage, compute, and security. Data Pipelines: Design, build, and maintain robust and scalable and automated data pipelines for batch and streaming data ingestion of data and processing (Data bricks workflow). Data Architecture and Modeling: Design and implement robust data models and architectures that align with business requirements and support efficient data processing, analysis, and reporting. Orchestration: Utilize workflow orchestration tools to automate data pipeline execution and dependency management. Monitoring and Alerting: Integrate monitoring and alerting mechanisms to track pipeline health, identify performance bottlenecks, and proactively address issues. Strong Agile principles: Utilize Agile development methodologies, actively participating in sprint planning, daily stand ups, sprint reviews, and retrospectives. Be flexible and adaptable to changing requirements and priorities throughout the project lifecycle. Unity Catalog Good to Have: Github Actiom Datagog exposure Data Quality: Implement data quality checks and balances throughout the data pipeline, including profiling, validation, and root cause analysis, to ensure data accuracy, completeness, and consistency. CI/CD: Implement continuous integration and continuous delivery (CI/CD) practices for automated testing and deployment of data pipelines.
What you’ll do
Design, build, and maintain scalable automated data pipelines for batch and streaming ingestion using Databricks. Implement robust data models and architectures while integrating monitoring and alerting mechanisms to ensure pipeline health.
Requirements
Requires proficiency in Python, PySpark, and PyTest with strong knowledge of OOP concepts and Azure cloud platforms. Experience with Delta Live Tables, Unity Catalog, and Agile methodologies is essential.
Listed skills
- Microsoft Azure · Preferred
- CI/CD · Preferred
- Agile · Preferred
- Python · Preferred
Other relevant skills
Identified from the job description. Confirm important requirements above.
- Python
- PySpark
- PyTest
- Databricks
- Delta Live Tables
- Azure
- OOP
- Data Pipelines
- Data Modeling
- Unity Catalog
- Agile
- CI/CD
- GitHub Actions
- Data Quality
- Orchestration
- Data Architecture
Job areas
- Data & Analytics
- Technology
- Software
- Engineering
- Consulting
