Senior Site Reliability Engineer
Top Benefits
About the role
Who you are
- This role demands a blend of deep technical expertise, a proactive mindset, and the ability to take vague or evolving challenges and refine them into robust, workable solutions
- 7+ years of experience in SRE, DevOps, or a related role
- Understanding of SRE best practices, architectures, and methods
- Good knowledge on resiliency patterns and cloud security
- Strong programming proficiency in Python, Golang, or Javascript
- Rust experience is advantageous
- Demonstrated experience with AWS and modern cloud architectures
- Proficiency in Helm, Terraform, and CI/CD tools like Github Actions and ArgoCD
- Hands-on experience with Kubernetes/EKS and GitOps methodologies
- Proven track record with monitoring tools such as Prometheus, OpenTelemetry, as well as familiarity with the LGTM stack, or other comparable tools
- Blockchain experience is advantageous, offering a unique perspective on distributed systems and security
- Exceptional problem-solving skills with a knack for translating vague requirements into clear, strategic plans
- Ability to engage in technical discussions and be part of the decision making process
- Strong problem-solving skills and capability to work on complex systems
- Experience in working within an Agile environment
- Experience in working with a distributed team
- Strong communication and collaboration abilities to work seamlessly across different teams
- A proactive and innovative mindset, with a passion for continuous improvement and operational excellence
What the job involves
- As a Senior SRE, you will be a key player in shaping the reliability and performance of our systems across our cloud infrastructure. You will design and implement solutions that improve our service reliability, automate routine tasks, and facilitate smooth collaboration between development and operations teams
- Infrastructure & Automation:
- Design, build, and maintain scalable and highly available systems, primarily on AWS, using best practices
- Manage and optimize Kubernetes clusters for high availability and performance, extending them when it makes sense to expand functionality
- Leverage GitOps principles to automate deployments and manage container orchestration
- Implement and manage CI/CD pipelines ensuring seamless, high-quality deployments, finding and removing bottlenecks, improving performance and working alongside teams to refine feedback loops and automate toil away
- Develop automation tools and scripts to improve operational efficiency
- Monitoring & Incident Response:
- Implement robust monitoring solutions with Prometheus and related tooling to ensure system health and performance
- Participate in on-call rotations and lead incident response efforts, turning challenges into learning opportunities
- Collaborate with dev teams to define and implement SLOs/SLIs
- Problem Solving & Communication:
- Take vague or loosely defined problems, work closely with cross-functional teams, and distill them into clear, actionable plans
- Communicate technical solutions and incident retrospectives effectively across both technical and non-technical stakeholders
- Innovation & Continuous Improvement:
- Evaluate and adopt new technologies, with a special advantage for candidates with blockchain experience, to keep our systems at the cutting edge
- Document processes and best practices, ensuring that knowledge is shared across the team and continuously improved
- Strive to strike a balance between effective delivery of goals and a measurable high standard of these goals. Always apply a layer of polish and due diligence when delivering
Benefits
- Flexible schedule
- Remote work - ability to work anywhere
- Laptop reimbursement
- New starter package to buy hardware essentials (headphones, monitor, etc)
- Learning & Development Opportunities
- Minimum 4 weeks of PTO + Sick Leave plan
- Medical, Dental, and Vision benefits coverage through Anthem with 100% premium cost covered by IO Global for the employee and dependents
- Health Savings Account
- Life Insurance
- Monthly Health Stipend to use towards any wellness or medical coverage/service
- Pension
About IO Global
Io Global is a leading ISP with a large customer base in Afghanistan specialising in broadband solutions with a high focus on Internet connectivity services. Io Global provides WiMAX, VSAT. Point to Point Wireless based Broadband Internet Access, to large Corporates, Banks, SMEs, SOHOs and Homes.
Io Global holds an Internet Service Provider License from the Ministry of Communications, Government of Afghanistan.
With several years of proven track-record of intensive on-field experience, Io Global are one of the foremost WiMax, VSAT/RF experts in Afghanistan. Founded by a team of professionals in the field of satellite communications, information technology and Telecoms, Io Global specialises in short and long-term turnkey contracts for services with operating engineering staffers, microwave, satellite and IP, design engineers and project managers.
With a satisfied client base including several multinational corporations, Io Global provides cost-efficient services with strategic business relationships with leading IT, microwave and Satellite services companies. Our service footprint covers the Indian sub-continent where we deploy a varied inventory of leased satellite space segment on selected satellites to deliver a cost-effective service into this geographic region.
The Company focuses on extreme customer satisfaction by implementing tried and tested systems and procedures, fast response time, quick turn ON time, proper documentation with access to state of the art test and measurement equipment and labs.
We have exclusive tie-ups with various companies all around the world, providing them with global support for all their Broadband Wireless products.
Senior Site Reliability Engineer
Top Benefits
About the role
Who you are
- This role demands a blend of deep technical expertise, a proactive mindset, and the ability to take vague or evolving challenges and refine them into robust, workable solutions
- 7+ years of experience in SRE, DevOps, or a related role
- Understanding of SRE best practices, architectures, and methods
- Good knowledge on resiliency patterns and cloud security
- Strong programming proficiency in Python, Golang, or Javascript
- Rust experience is advantageous
- Demonstrated experience with AWS and modern cloud architectures
- Proficiency in Helm, Terraform, and CI/CD tools like Github Actions and ArgoCD
- Hands-on experience with Kubernetes/EKS and GitOps methodologies
- Proven track record with monitoring tools such as Prometheus, OpenTelemetry, as well as familiarity with the LGTM stack, or other comparable tools
- Blockchain experience is advantageous, offering a unique perspective on distributed systems and security
- Exceptional problem-solving skills with a knack for translating vague requirements into clear, strategic plans
- Ability to engage in technical discussions and be part of the decision making process
- Strong problem-solving skills and capability to work on complex systems
- Experience in working within an Agile environment
- Experience in working with a distributed team
- Strong communication and collaboration abilities to work seamlessly across different teams
- A proactive and innovative mindset, with a passion for continuous improvement and operational excellence
What the job involves
- As a Senior SRE, you will be a key player in shaping the reliability and performance of our systems across our cloud infrastructure. You will design and implement solutions that improve our service reliability, automate routine tasks, and facilitate smooth collaboration between development and operations teams
- Infrastructure & Automation:
- Design, build, and maintain scalable and highly available systems, primarily on AWS, using best practices
- Manage and optimize Kubernetes clusters for high availability and performance, extending them when it makes sense to expand functionality
- Leverage GitOps principles to automate deployments and manage container orchestration
- Implement and manage CI/CD pipelines ensuring seamless, high-quality deployments, finding and removing bottlenecks, improving performance and working alongside teams to refine feedback loops and automate toil away
- Develop automation tools and scripts to improve operational efficiency
- Monitoring & Incident Response:
- Implement robust monitoring solutions with Prometheus and related tooling to ensure system health and performance
- Participate in on-call rotations and lead incident response efforts, turning challenges into learning opportunities
- Collaborate with dev teams to define and implement SLOs/SLIs
- Problem Solving & Communication:
- Take vague or loosely defined problems, work closely with cross-functional teams, and distill them into clear, actionable plans
- Communicate technical solutions and incident retrospectives effectively across both technical and non-technical stakeholders
- Innovation & Continuous Improvement:
- Evaluate and adopt new technologies, with a special advantage for candidates with blockchain experience, to keep our systems at the cutting edge
- Document processes and best practices, ensuring that knowledge is shared across the team and continuously improved
- Strive to strike a balance between effective delivery of goals and a measurable high standard of these goals. Always apply a layer of polish and due diligence when delivering
Benefits
- Flexible schedule
- Remote work - ability to work anywhere
- Laptop reimbursement
- New starter package to buy hardware essentials (headphones, monitor, etc)
- Learning & Development Opportunities
- Minimum 4 weeks of PTO + Sick Leave plan
- Medical, Dental, and Vision benefits coverage through Anthem with 100% premium cost covered by IO Global for the employee and dependents
- Health Savings Account
- Life Insurance
- Monthly Health Stipend to use towards any wellness or medical coverage/service
- Pension
About IO Global
Io Global is a leading ISP with a large customer base in Afghanistan specialising in broadband solutions with a high focus on Internet connectivity services. Io Global provides WiMAX, VSAT. Point to Point Wireless based Broadband Internet Access, to large Corporates, Banks, SMEs, SOHOs and Homes.
Io Global holds an Internet Service Provider License from the Ministry of Communications, Government of Afghanistan.
With several years of proven track-record of intensive on-field experience, Io Global are one of the foremost WiMax, VSAT/RF experts in Afghanistan. Founded by a team of professionals in the field of satellite communications, information technology and Telecoms, Io Global specialises in short and long-term turnkey contracts for services with operating engineering staffers, microwave, satellite and IP, design engineers and project managers.
With a satisfied client base including several multinational corporations, Io Global provides cost-efficient services with strategic business relationships with leading IT, microwave and Satellite services companies. Our service footprint covers the Indian sub-continent where we deploy a varied inventory of leased satellite space segment on selected satellites to deliver a cost-effective service into this geographic region.
The Company focuses on extreme customer satisfaction by implementing tried and tested systems and procedures, fast response time, quick turn ON time, proper documentation with access to state of the art test and measurement equipment and labs.
We have exclusive tie-ups with various companies all around the world, providing them with global support for all their Broadband Wireless products.