Back to job search
A
AmpstekVerified Job Source

Data Engineer

The role involves architecting ETL pipelines and managing vector databases to provide clean, real-time data for AI systems. Responsibilities include building event-driven architectures to continuously refresh embeddings and indexes while enforcing data governance.

  • On-site
  • Toronto, ON
  • Posted Aug 6, 2026
  • Apply by Sep 5, 2026
  • 1 position

Job summary

Job Title: Data Engineer Job Location: Toronto, Brampton, Canada The data backbone owner who ensures our AI systems have clean, structured, real-time data to reason over. About the Role You will architect ETL pipelines, manage vector databases, enforce governance, and build real-time data flows that continuously update embeddings and indexes. Your work ensures our AI agents operate with fresh, trustworthy information. What You Will Do • Data Pipelines: Build ETL flows for structured/unstructured data, ensuring normalization, deduplication, and semantic consistency. • Vector Infrastructure: Manage pgvector, Azure AI Search, Redis vector indexing, and hybrid search layers. • Data Governance: Implement zero-trust access, privacy controls, and compliance within AI context pipelines. • Real-time Processing: Build event-driven architectures that continuously refresh embeddings and indexes. Required Qualifications • Deep experience with distributed data systems, SQL, and orchestration tools. • Experience tuning high-throughput database infrastructure. • Knowledge of Google GECX is a plus. • Familiarity with chunking strategies and embedding models Thanks, and Regards Anil Kumar | Technical Recruiter Email: anil.k@ampstek.com Desk: 6095361083 www.ampstek.com

What you’ll do

The role involves architecting ETL pipelines and managing vector databases to provide clean, real-time data for AI systems. Responsibilities include building event-driven architectures to continuously refresh embeddings and indexes while enforcing data governance.

Requirements

Candidates need deep experience with distributed data systems, SQL, and high-throughput database infrastructure. Knowledge of chunking strategies, embedding models, and Google GECX is also preferred.

Listed skills

  • SQLPreferred

Other relevant skills

Identified from the job description. Confirm important requirements above.

  • ETL Pipelines
  • Vector Databases
  • Data Governance
  • Real-time Processing
  • SQL
  • Distributed Data Systems
  • Pgvector
  • Azure AI Search
  • Redis Vector Indexing
  • Hybrid Search
  • Chunking Strategies
  • Embedding Models
  • Orchestration Tools

Job areas

  • Data & Analytics
  • Technology
  • Software
  • Engineering
  • Consulting

Additional details

Minimum experience
5+ years
Apply by
Sep 5, 2026
Posting language
English
Working hours
40 hours per week
Seniority
Mid-Senior level
Application method
Direct apply is available