Jobs.ca
Jobs.ca
Language
Syndesus logo

Sr Data Engineer

Syndesus1 day ago
Canada
Senior Level
Full-Time

About the role

Senior Data Engineer

Healthtech SaaS - ML Platform Team - Remote (Canada)

About the Company

Our client is a well-funded healthtech company serving over 140,000 independent medical practices across the United States. Formed through the merger of two established healthcare software businesses - they now operate as the full-stack platform for independent care providers. T

The Role

This is the first dedicated Data Engineer hire at the company - a greenfield mandate with real influence over how the data foundation gets built. You will sit on the ML team, working directly alongside Machine Learning Engineers to design and own the data infrastructure that powers model training, inference, and monitoring at scale. The team today has strong software engineers doing DE work, but without the rigor: no data lineage, no pipeline observability, no quality guarantees. You come in and change that.

The data landscape is a mix of structured claims and billing data alongside unstructured clinical notes - terabytes of data, billions of rows, processed across a Medallion lakehouse architecture on GCP. You will own the layout of that architecture, build the pipelines that feed ML training runs, and establish the tooling and standards the team operates from going forward.

What You'll Own

Data lake architecture - take a partially-built Medallion structure (bronze/silver/gold) and lay it out properly for long-term ML use; make decisions that will be hard to undo, so make them well. ML training pipelines - extract from siloed data stores, process and transform, feed structured and unstructured data into the training pipeline; understand at a high level how training works so you can supply the right datasets. Feature store design and ownership - build and own both offline and online feature stores; understand the distinction and the tradeoffs between them in production. Pipeline observability and data quality - instrument pipelines with monitoring, alerting, and validation from day one; track lineage, catch row drops, flag accuracy issues before downstream teams are affected. Kafka integration - collaborate with Java engineering teams to consume event-based data sources; understand how application events are published and how to pull them into the data pipeline reliably. Data quality frameworks - build automated checks and schema validation tooling that give the ML team confidence in what they are training on. Engineering design - lead architecture decisions, document trade-offs, and build systems others can extend; this is a software engineering role as much as a data role.

What You Bring

5+ years of software development experience, with 3+ years focused on Data Engineering supporting data science or ML teams. Proven experience building ML training pipelines and data lake / lakehouse architecture - not just analytics pipelines; you understand what it means to supply data to a model, not just a dashboard. Deep Python and SQL proficiency - production-grade code, not scripts; you write software that other engineers can maintain and extend. Strong Spark background - optimization at scale (partitioning, broadcast joins, predicate pushdown) for large dataset processing; terabyte-range workloads are the norm here. Hands-on experience with cloud data warehouse and lakehouse tooling - Snowflake, Databricks, Delta Lake, or equivalents; concepts transfer, specific tools do not need to match exactly. Feature store experience - offline vs online serving, freshness guarantees, training-serving skew prevention; ideally Feast, Tecton, or similar. MLOps fundamentals - data lineage, versioning, quality monitoring, data leakage prevention; you treat these as table stakes, not extras. Kafka or event-driven data source experience - consuming from message queues or pub/sub systems as part of an ingestion pipeline. Airflow or equivalent orchestration at production scale - set up from scratch, not inherited. CI/CD, monitoring, and alerting as standard practice - you deploy data systems, not just build them.

Nice to Have

Experience with unstructured data pipelines - clinical notes, documents, or similar text-heavy unstructured sources. Background in healthcare software or working with HL7, FHIR, or structured clinical/claims data. RAG pipelines or vector search infrastructure in production. Data versioning tooling (DVC, LakeFS) in a production ML context. GCP experience - Vertex AI, BigQuery, Cloud Composer, Pub/Sub; the team runs on GCP. Published research or open-source contributions in data engineering, distributed systems, or machine learning.

About Syndesus

Staffing and Recruiting
11-50
Founded in 2014

Syndesus builds engineering teams in Canada for VC-backed startups in the U.S., and offers Professional Employer Organization (PEO) services for U.S. companies seeking to employ workers remotely in Canada.

Additionally, Syndesus can assist foreign-born tech workers (and their U.S. employers) with options for working remotely in Canada if they cannot stay in the U.S. due to immigration/work visa issues.

Learn more at syndesus.com

Similar Jobs