AVISHEK BHANDARI
Open to opportunitiesData Engineer | SQL, Python, PySpark | AWS & Azure Cloud Data Platforms | ETL/ELT | Data Warehousing
Clawson, Mi
About
Data Engineer with 6+ years of progressive experience designing, building, and operating secure, scalable ETL/ELT pipelines, cloud data workflows, data lake and warehouse solutions, and analytics-ready datasets across AWS and Azure environments. Strong hands-on background in SQL, Python, PySpark, Spark SQL, AWS Glue, Azure Data Factory, Databricks, Amazon Redshift, Snowflake, ADLS Gen2, and S3. Experienced in ingesting, cleaning, validating, transforming, and modeling structured and semi-structured data for healthcare, financial services, reporting, reconciliation, and analytics use cases. Focused on building reliable, secure, well-documented data pipelines that improve data quality, support reconciliation, reduce reporting issues, and deliver business-ready datasets for analytics, reporting, and operational decision-making.
Skills
- Agile
- Amazon Web Services
- CI/CD
- Communication
- Cooking
- Data Validation
- Docker
- English
- Git
- GitHub
- HTML
- JavaScript
- Jira
- Kubernetes
- Laravel
- Leadership
- Microsoft Azure
- Microsoft Word
- MongoDB
- MySQL
- Node.js
- PHP
- PostgreSQL
- Power BI
- Problem solving
- Python
- SQL
- Tableau
- Tailwind CSS
- Teamwork
- Terraform
- Vue.js
Experience
Data Engineer
Mastercard
Dec 2024 to Present
Purchase, NY
Environment: AWS S3, AWS Glue, Glue Data Catalog, Amazon Redshift, Athena, Airflow, PySpark, Spark SQL, Python, SQL, dbt, Snowflake, IAM, KMS, CloudWatch, Parquet, Git, CI/CD • Designed and maintained AWS-based ETL pipelines using AWS Glue, PySpark, and SQL to process high-volume payment, transaction, settlement, and reconciliation datasets into an S3-based data lake. • Built Bronze, Silver, and Gold data layers to support downstream reporting, fraud analytics, chargeback analysis, settlement workflows, and business reconciliation use cases. • Developed SQL transformations and reporting-ready warehouse tables in Amazon Redshift and Snowflake, applying distribution keys, sort keys, reusable views, and query optimization techniques to improve analytics performance. • Developed and maintained dbt SQL models for staging, intermediate, and mart layers in Redshift/Snowflake, supporting standardized transformations, reusable business logic, data testing, documentation, and analytics-ready reporting tables. • Orchestrated batch workflows using Airflow and AWS-native services, including scheduled ingestion, transformation dependencies, retries, and failure handling. • Investigated pipeline failures, performed root-cause analysis, and improved retry, alerting, and controlled reprocessing procedures to strengthen production reliability. • Structured logging, data-quality checks, and reconciliation processes for nulls, duplicates, schema changes, row counts, rejected records, and source-to-target validation to improve pipeline reliability and reduce recurring reporting issues. • Supported secure handling of sensitive financial data using IAM least-privilege access, KMS encryption, audit logging, and controlled access patterns. • Tuned PySpark and Glue jobs using partitioning, optimized file formats, broadcast joins, caching, and Parquet-based storage patterns. • Partnered with analytics, reconciliation, fraud, and reporting teams to translate payment-data requirements into curated warehouse tables, reusable SQL models, and validated datasets for operational decision-making.
Data Engineer
UnitedHealth Group
Jan 2022 to Dec 2023
Minnetonka, MN
Environment: Azure Data Factory, Azure Databricks, Azure Synapse Analytics, ADLS Gen2, Azure SQL Database, Delta Lake, Azure Key Vault, Azure DevOps, PySpark, Spark SQL, Python, SQL, Power BI, RBAC, Git, CI/CD • Built and maintained Azure Data Factory pipelines using linked services, datasets, parameters, triggers, and validation activities to ingest healthcare data from databases, APIs, SFTP locations, and partner feeds. • Developed Azure Databricks notebooks using PySpark, Spark SQL, and Delta Lake to clean, standardize, deduplicate, and curate healthcare datasets across Bronze, Silver, and Gold layers in ADLS Gen2. • Created SQL-based transformation logic and dimensional data models in Azure Synapse and Azure SQL Database to support clinical, claims, eligibility, quality-measure, and revenue-cycle reporting. • Implemented data quality checks for schema validation, required fields, duplicate records, row counts, and reconciliation between source and target tables. • Supported PHI/PII protection patterns using Azure Key Vault, RBAC, encryption, audit logging, and secure access controls across pipeline and reporting layers. • Collaborated with clinical, finance, analytics, and reporting stakeholders to deliver curated Power BI-ready datasets, improving trust in operational dashboards, executive reporting, and healthcare data analysis. • Supported Spark and Delta Lake performance improvements through partitioning, file compaction, and query optimization patterns for large healthcare batch workloads. • Supported scheduled ADF and Databricks workloads by investigating failed pipeline runs, validating downstream data, coordinating controlled reprocessing, and documenting recurring operational issues. • Used Git and Azure DevOps to manage version control for ADF artifacts, Databricks notebooks, and SQL scripts across development and deployment workflows.
Data Engineer
Fusemachines
Jun 2019 to Dec 2021
New York, NY
Environment: Python, SQL, PySpark, Apache Spark, Airflow, Snowflake, PostgreSQL, MySQL, Parquet, Docker, Git, Agile/Scrum • Supported senior data engineers in developing and maintaining ETL/ELT pipelines using Python, SQL, and PySpark to ingest structured and semi-structured data from databases, APIs, and flat files. • Wrote SQL queries for data extraction, profiling, validation, duplicate checks, null checks, and source-to-target comparison to improve dataset reliability. • Assisted senior engineers in building Spark-based batch transformation jobs for data cleaning, standardization, joining, aggregation, and preparation of analytics-ready datasets. • Created and maintained staging tables, SQL views, and curated datasets in PostgreSQL, MySQL, and Snowflake to support analytics and data science teams. • Helped develop Airflow DAGs for scheduled pipeline execution, dependency handling, retries, and basic failure monitoring. • Collaborated with analysts, data scientists, and senior data engineers to understand source data, document transformation logic, and deliver clean datasets for reporting and model development. • Used Git, Docker-based development environments, and Agile practices to version code, test pipeline changes, and support sprint-based delivery. • Expanded responsibilities over time from SQL-based data preparation and validation to supporting PySpark transformations, Airflow orchestration, and Snowflake data-delivery workflows.
Education
Trine University
Master of Science, Information Studies
USA
Itahari International College
Bachelor's, Information Technology
Nepal
Licences & certifications
Microsoft Certified: Azure Fundamentals (AZ-900)
AWS Certified Cloud Practitioner
