Back to job search
C
ClifyXVerified Job Source

Data Engineer

Architect and manage ETL pipelines and vector indexing infrastructure to support AI systems and RAG memory. Ensure data governance, compliance, and real-time data flows to provide AI agents with trustworthy information.

  • On-site
  • Brampton, ON
  • Posted Jul 31, 2026
  • Apply by Aug 30, 2026
  • 1 position

Job summary

Position Title: Data Engineer Location: Brampton Canada Duration: 12 Months Contract • 5-7 yr of exp ETL & Data Pipelines • Vector Databases (pgVector, Azure AI Search, Redis) • Real-time Data Processing • Data Governance & Compliance The data backbone owner who ensures AI systems have clean, structured, real-time data to reason over. You will design ingestion pipelines, vector indexing infrastructure, and governance layers that keep RAG and agent memory systems accurate and compliant. About the Role You will architect ETL pipelines, manage vector databases, enforce governance, and build real-time data flows that continuously update embeddings and indexes. Your work ensures AI agents operate with fresh, trustworthy information. What You Will Do · Data Pipelines: Build ETL flows for structured/unstructured data, ensuring normalization, deduplication, and semantic consistency. · Vector Infrastructure: Manage pgvector, Azure AI Search, Redis vector indexing, and hybrid search layers. · Data Governance: Implement zero-trust access, privacy controls, and compliance within AI context pipelines. · Real-time Processing: Build event-driven architectures that continuously refresh embeddings and indexes. Required Qualifications · Deep experience with distributed data systems, SQL, and orchestration tools. · Experience tuning high-throughput database infrastructure. · Knowledge of Google’s GECX is a plus. · Familiarity with chunking strategies and embedding models. Skillset Requirements · ETL & Data Modeling: Designing pipelines for structured/unstructured data, normalization, deduplication, and semantic consistency. · Vector Databases: pgvector, Redis, Azure AI Search, hybrid search, and index optimization. · Distributed Data Systems: Kafka, Spark, Flink, or similar event-driven architectures. · Data Governance: Zero-trust access, privacy controls, compliance, and auditability. · Real-time Embedding Updates: Event-driven refresh pipelines for RAG and agent memory systems. · Chunking & Embeddings: Semantic chunking, metadata tagging, and embedding model selection. · Search Infrastructure: BM25, hybrid search, inverted indexes, and ranking algorithms. · Performance Tuning: High-throughput read/write optimization. · Data Quality & Lineage: Validation, schema enforcement, and lineage tracking (e.g., Great Expectations, OpenLineage). Regards, Hasan Choudhary (Executive Recruiter ) Tel – 908-279-1243 Fax - 732-909-2631 Email - hasan@ClifyX.com

What you’ll do

Architect and manage ETL pipelines and vector indexing infrastructure to support AI systems and RAG memory. Ensure data governance, compliance, and real-time data flows to provide AI agents with trustworthy information.

Requirements

Requires 5-7 years of experience in ETL, distributed data systems, and vector databases. Candidates should be familiar with chunking strategies, embedding models, and high-throughput database tuning.

Other relevant skills

Identified from the job description. Confirm important requirements above.

  • ETL
  • Data Pipelines
  • Vector Databases
  • pgVector
  • Azure AI Search
  • Redis
  • Real-time Data Processing
  • Data Governance
  • SQL
  • Kafka
  • Spark
  • Flink
  • RAG
  • Semantic Chunking
  • Embedding Models
  • Hybrid Search

Job areas

  • Data & Analytics
  • Technology
  • Software
  • Engineering
  • Consulting

Additional details

Minimum experience
5+ years
Apply by
Aug 30, 2026
Posting language
English
Working hours
40 hours per week
Seniority
Mid-Senior level
Application method
Direct apply is available