Data Engineer
Lead the design and development of new PySpark data engineering capabilities and feature-engineering pipelines to support machine learning initiatives. This includes reverse-engineering legacy SAS programs to translate business logic into scalable PySpark code.
- Hybrid
- Toronto, ON
- Posted Aug 5, 2026
- Apply by Sep 4, 2026
- 1 position
Job summary
Iris's Fortune 100 direct client is looking for a Data Engineer in Toronto, ON. Please find the job description below: Job Title: Data Engineer Location: Toronto, ON (Hybrid – 3 days a week in office) Mode of interview: Final Round In-person Interview Domain: Banking / Financial We are seeking a highly skilled Data Engineer to help build a net-new PySpark data engineering capability from the ground up. The initial focus of this role will be on designing and developing net-new feature-engineering pipelines to support downstream data science initiatives. Key Responsibilities Foundation & Feature Engineering: Lead the foundational setup of Python/PySpark development environments, actively managing complex package dependencies and virtual environments using Conda/Anaconda. Pipeline Development: Design, build, and deploy net-new data pipelines focused heavily on feature engineering and data preparation to feed downstream machine learning and analytics use cases. Code Translation: Analyze and reverse-engineer existing SAS programs (developed by business users and data scientists) to accurately extract business rules, data transformations, and calculations. Platform Modernization: Translate extracted SAS logic into efficient, scalable, and functionally equivalent PySpark code. Performance Tuning: Optimize PySpark code for performance, scalability, and maintainability within distributed data processing environments. Stakeholder Collaboration: Collaborate closely with business users and data scientists to clarify requirements, validate feature outputs, and resolve discrepancies during the migration process. Required Qualifications 8+ years of hands-on experience in data engineering, data pipeline development, or building data platforms. Strong programming expertise in Python and PySpark for distributed data processing. Proven experience building data pipelines specifically for feature engineering, data curation, and advanced data preparation. Ability to read, interpret, and reverse-engineer legacy SAS code (such as SAS data steps, procedures, and macros) to extract complex business logic. (Note: Deep SAS development experience is a plus, but the ability to translate it is the core requirement). Experience establishing and managing Python environments, ensuring reproducibility, and handling library dependencies using Anaconda/Conda. Advanced SQL skills and experience working with large-scale structured datasets. Experience with code migration, platform modernization, or translating legacy codebases into modern open-source stacks. Strong analytical, problem-solving, and communication skills, with a track record of successfully interfacing directly with business stakeholders. About Iris Software Inc. With 4,000+ associates and offices in India, the U.S.A., and Canada, Iris Software delivers technology services and solutions that help clients complete fast, far reaching digital transformations and achieve their business goals. A strategic partner to Fortune 500 and other top companies in financial services and many other industries, Iris provides a value-driven approach - a unique blend of highly-skilled specialists, software engineering expertise, cutting-edge technology, and flexible engagement models. High customer satisfaction has translated into long-standing relationships and preferred-partner status with many of our clients, who rely on our 30+ years of technical and domain expertise to future-proof their enterprises. Associates of Iris work on mission-critical applications supported by a workplace culture that has won numerous awards in the last few years, including Certified Great Place to Work in India; Top 25 GPW in IT & IT-BPM; Ambition Box Best Place to Work, #3 in IT/ITES; and Top Workplace NJ-USA.
What you’ll do
Lead the design and development of new PySpark data engineering capabilities and feature-engineering pipelines to support machine learning initiatives. This includes reverse-engineering legacy SAS programs to translate business logic into scalable PySpark code.
Requirements
Requires over 8 years of experience in data engineering with strong expertise in Python, PySpark, and SQL. Candidates must be able to interpret legacy SAS code and manage Python environments using Conda/Anaconda.
Listed skills
- SQLPreferred
- PythonPreferred
Other relevant skills
Identified from the job description. Confirm important requirements above.
- Python
- PySpark
- SAS
- Feature Engineering
- Data Pipeline Development
- Conda
- Anaconda
- SQL
- Code Migration
- Platform Modernization
- Distributed Data Processing
- Data Curation
- Reverse Engineering
- Performance Tuning
- Stakeholder Collaboration
- Data Preparation
Job areas
- Data & Analytics
- Technology
- Engineering
- Software
- Finance & Accounting
Additional details
- Minimum experience
- 10+ years
- Apply by
- Sep 4, 2026
- Posting language
- English
- Working hours
- 40 hours per week
- Office presence
- 3 days per week
- Seniority
- Mid-Senior level
- Application method
- Direct apply is available
