Data Engineer
The role involves architecting ETL pipelines and managing vector databases to provide clean, real-time data for AI systems. Responsibilities include building event-driven architectures to continuously refresh embeddings and indexes while enforcing data governance.
- On-site
- Toronto, ON
- Posted Aug 6, 2026
- Apply by Sep 5, 2026
- 1 position
Job summary
Job Title: Data Engineer Job Location: Toronto, Brampton, Canada The data backbone owner who ensures our AI systems have clean, structured, real-time data to reason over. About the Role You will architect ETL pipelines, manage vector databases, enforce governance, and build real-time data flows that continuously update embeddings and indexes. Your work ensures our AI agents operate with fresh, trustworthy information. What You Will Do • Data Pipelines: Build ETL flows for structured/unstructured data, ensuring normalization, deduplication, and semantic consistency. • Vector Infrastructure: Manage pgvector, Azure AI Search, Redis vector indexing, and hybrid search layers. • Data Governance: Implement zero-trust access, privacy controls, and compliance within AI context pipelines. • Real-time Processing: Build event-driven architectures that continuously refresh embeddings and indexes. Required Qualifications • Deep experience with distributed data systems, SQL, and orchestration tools. • Experience tuning high-throughput database infrastructure. • Knowledge of Google GECX is a plus. • Familiarity with chunking strategies and embedding models Thanks, and Regards Anil Kumar | Technical Recruiter Email: anil.k@ampstek.com Desk: 6095361083 www.ampstek.com
What you’ll do
The role involves architecting ETL pipelines and managing vector databases to provide clean, real-time data for AI systems. Responsibilities include building event-driven architectures to continuously refresh embeddings and indexes while enforcing data governance.
Requirements
Candidates need deep experience with distributed data systems, SQL, and high-throughput database infrastructure. Knowledge of chunking strategies, embedding models, and Google GECX is also preferred.
Listed skills
- SQLPreferred
Other relevant skills
Identified from the job description. Confirm important requirements above.
- ETL Pipelines
- Vector Databases
- Data Governance
- Real-time Processing
- SQL
- Distributed Data Systems
- Pgvector
- Azure AI Search
- Redis Vector Indexing
- Hybrid Search
- Chunking Strategies
- Embedding Models
- Orchestration Tools
Job areas
- Data & Analytics
- Technology
- Software
- Engineering
- Consulting
Additional details
- Minimum experience
- 5+ years
- Apply by
- Sep 5, 2026
- Posting language
- English
- Working hours
- 40 hours per week
- Seniority
- Mid-Senior level
- Application method
- Direct apply is available
