HG

Hardik Gandhi

Open to opportunities

Site Reliability Engineer|DevOps Engineer

Ottawa, ON

Sign in to follow
Stripe
Carleton University

About

Site Reliability Engineer with 3+ years of experience building highly available, scalable, and secure cloud platforms across fintech and SaaS environments using AWS, Kubernetes, Terraform, Docker, Linux, and Infrastructure as Code. Proven expertise in SRE, DevOps, AI-assisted operations, CI/CD automation, observability, incident response, cloud reliability, and distributed systems, delivering resilient production platforms that support millions of users and high-volume transactions. Experienced in improving system reliability, automating infrastructure, optimizing cloud performance, strengthening security, and enabling faster software delivery through GitOps, DevSecOps, and modern platform engineering practices. Strong collaborator who partners with engineering, product, and security teams to solve complex operational challenges, reduce production risk, and build reliable, cost-efficient, and scalable cloud infrastructure that accelerates business growth.

Skills

  • Agile
  • Docker
  • Git
  • Kubernetes
  • Linux
  • Python
  • Redis
  • Scrum
  • Terraform

Experience

  1. Site Reliability Engineer

    Stripe

    Apr 2025 to Present

    Canada

    • Built AI-assisted incident response workflows with Python, OpenAI APIs, Kubernetes, AWS, and Grafana that automate log correlation, runbook suggestions, and root cause analysis, cutting average resolution time by 42 minutes across 180+ production incidents a year for global platform engineering teams. • Built highly available Kubernetes infrastructure on AWS EKS using Terraform, Helm, ArgoCD, and GitHub Actions to standardize multi-region deployments, supporting 520+ microservices that process nearly 24 million payment transactions daily on Stripe's global merchant platform. • Strengthened platform observability with Prometheus, Grafana, OpenTelemetry, Loki, and CloudWatch dashboards to set up SLI/SLO monitoring, cutting critical alert noise by roughly 1,400 alerts a month and giving SRE and engineering teams clearer operational visibility. • Automated infrastructure provisioning with Terraform modules, AWS CloudFormation, IAM, and policy-as-code, eliminating repetitive manual work and cutting environment provisioning time from 5 hours to under 20 minutes for development and platform teams. • Improved production reliability by designing progressive delivery pipelines with Argo Rollouts, automated canary deployments, health validation, and rollback automation, letting engineering teams safely ship 350+ production releases a month with no customer-facing disruptions. • Optimized Kubernetes resource utilization by tuning Horizontal Pod Autoscaler, Cluster Autoscaler, Redis caching, and container resource allocation, lowering compute consumption by nearly 1,900 vCPUs a month while maintaining 99% availability during global payment traffic peaks.

  2. DevOps Engineer

    Zoho

    Jul 2021 to Jul 2023

    India

    • Designed and built enterprise CI/CD pipelines with Jenkins, GitLab CI, Maven, Docker, Kubernetes, and Ansible to automate application delivery, cutting release cycles from 6 hours to about 30 minutes across 85+ production applications for multiple Zoho SaaS products. • Provisioned secure AWS infrastructure with Terraform, Infrastructure as Code, and reusable deployment templates to standardize cloud environments, supporting 260+ EC2 instances, load balancers, and production databases across engineering teams. • Built centralized monitoring and alerting platforms with Prometheus, Grafana, ELK Stack, and CloudWatch, improving infrastructure visibility and reducing high-priority production incidents by 110 a year through proactive health monitoring and alert tuning. • Automated server configuration management, patching, deployments, and operational workflows using Ansible, Python, and Bash, eliminating repetitive manual tasks and improving deployment consistency for DevOps and release engineering teams. • Modernized container orchestration by deploying Kubernetes clusters with Helm, rolling deployments, readiness probes, and autoscaling, supporting 240+ containerized microservices with reliable scalability across production environments. • Integrated SonarQube, Trivy, dependency scanning, and secret detection into CI/CD pipelines, catching and fixing 950+ security and code quality issues before production deployment, improving release quality for development teams.

Education

  1. Carleton University

    Master of Engineering, Computer Engineering

    Ottawa, ON

    2023 to 2025

  2. Charotar University of Science and Technology.

    Bachelor of Technology, Information Technology

    Anand, Gujarat, India

    2019 to 2023

Licences & certifications

  • Azure Administrator Associate (AZ-104)

  • AWS Certified Cloud Practitioner

  • Oracle Cloud AI Foundations Associate