Back to job search
VA
Veeda AIVerified Job Source

Machine Learning Engineer - Data Curation

Build and maintain high-throughput multimodal data pipelines for cleaning, filtering, and augmenting image and video datasets. Translate product goals into data strategies and coordinate large-scale labeling efforts to support foundation world models.

  • Hybrid
  • Toronto, ON
  • Posted Jul 21, 2026
  • 1 position

Job summary

About Us Veeda AI is building the next generation of multimodal foundation world models for Physical AI. We're a small, fast-moving team of engineers and researchers from leading AI labs, tackling some of the most challenging problems at the intersection of AI, robotics, and embodied intelligence. If you're excited about pushing the boundaries of what's possible with Physical AI, you'll have the opportunity to make an outsized impact from day one. Responsibilities Data Processing at Scale: Build and maintain high-throughput image and video data pipelines—cleaning, filtering, augmenting, and transforming multimodal datasets Data Strategy from Product Goals: Translate product objectives and training strategies into concrete data processing plans, defining required dataset characteristics, sourcing strategies, and preparation protocols for model consumption. Iterative Quality Evaluation: Evaluate processed datasets against rigorous quality benchmarks, diagnose data-side failure modes, and iterate on processing strategies until datasets meet the high bar our foundation models demand. Large-Scale Labeling Coordination: Coordinate and drive data labeling efforts with annotation teams, establishing labeling guidelines, reviewing outputs, and ensuring consistency across large-scale annotation campaigns. Requirements You have strong Python programming skills and write clean, modular, production-grade data processing code. You have hands-on experience with ML frameworks such as PyTorch—understanding model data requirements, tensor formats, and training data flow well enough to prepare data that researchers can consume directly. You hold exceptionally high standards for data quality and can assess data reliability, accuracy, and coverage systematically. You are familiar with Computer Vision domain concepts and understand how image and video data characteristics impact downstream generative model performance. You communicate clearly, document your work thoroughly, and collaborate effectively with researchers, engineers, and labeling partners across time zones. Production Quality, Agent Velocity: Your daily workflow runs through AI coding harnesses (e.g., AI agents/assistants), without sacrificing software engineering rigor. You review agent code diffs with the same scrutiny as a team member's PR, recognize AI code generation failure modes, and ship rapidly without introducing technical debt or "slop." Nice to Have Experience with Computer Vision tasks related to image or video generation model training (e.g., diffusion models, autoregressive transformers, GANs). Fluency with annotation formats such as COCO, PASCAL VOC, or custom labeling schemas. Hands-on experience with data orchestration frameworks (e.g., Airflow, Dagster, Prefect, Luigi). Experience with distributed data processing systems (e.g., Ray, Spark, Dask).

What you’ll do

Build and maintain high-throughput multimodal data pipelines for cleaning, filtering, and augmenting image and video datasets. Translate product goals into data strategies and coordinate large-scale labeling efforts to support foundation world models.

Requirements

Requires strong Python skills, experience with PyTorch, and a deep understanding of Computer Vision and data quality standards. Candidates should be proficient in using AI coding tools to maintain high software engineering rigor and velocity.

Other relevant skills

Identified from the job description. Confirm important requirements above.

  • Python
  • PyTorch
  • Computer Vision
  • Data Pipeline Construction
  • Multimodal Datasets
  • Data Labeling Coordination
  • AI Coding Assistants
  • Data Quality Evaluation
  • Image Processing
  • Video Processing
  • Distributed Data Processing
  • Data Orchestration
  • Dask (Software)
  • Transformer (Machine Learning Model)
  • Workflow Management
  • Apache Airflow
  • Generative Adversarial Networks
  • Data Labeling
  • Data Strategy
  • AI Agents
  • Research
  • Artificial Intelligence
  • Autoregressive Model
  • Code Generation
  • Communication
  • Data Processing
  • Dataflow
  • Data Modeling
  • Data Quality
  • Distributed Data Store
  • Failure Causes
  • Packaging And Labeling
  • Python (Programming Language)
  • Pascal (Programming Language)
  • Robotics
  • Software Engineering
  • Coordinating
  • Reliability
  • Technical Debt
  • Luigi (Python Package)
  • Machine Learning Frameworks
  • Dagster

Job areas

  • Technology
  • Software
  • Data & Analytics
  • Engineering
  • Science & Research
  • Machine Learning Data Engineer
  • Machine Learning Engineer
  • Software Developers
  • Computer and Information Research Scientists

Additional details

Minimum experience
2+ years
Posting language
English
Working hours
40 hours per week
Location requirements
Country, Canada