S
Site Reliability Engineer
Dubai, United Arab Emirates · Full Time
Be the first to apply
- Experience
- 4+ yrs
- Salary
- —
- Openings
- 1
- Posted
- 1 ವಾರ ಹಿಂದೆ
- Work mode
- In office
- Resume
- Required to apply
Where you'll work
Sign in to tell us what does and doesn't work for you here — it sharpens every match we show you.
Job description
Role Overview
This position merges software engineering insights with infrastructure management, collaborating closely with product and platform teams to ensure systems remain observable, scalable, and reliable. You will significantly influence how BOT operates and develops software at scale, supporting over 500 personnel and expanding.
Primary Responsibilities
- Manage the reliability of essential services by defining and monitoring SLIs, SLOs, and SLAs; lead post-mortem analyses without blame and work to reduce mean time to recovery through effective incident handling.
- Develop and maintain AWS cloud infrastructure primarily using Terraform, ensuring that all environments are reproducible, version-controlled, and auditable.
- Optimize and scale Kubernetes container platforms, including capacity planning, resource tuning, autoscaling, and managing the full lifecycle of clusters.
- Enhance end-to-end observability by instrumenting services with Prometheus and Grafana, creating actionable alerts and minimizing alert noise by continuous refinement.
- Accelerate continuous integration and deployment pipelines using GitHub Actions or GitLab CI to facilitate frequent, secure deployments, embedding reliability checks within delivery workflows.
- Collaborate with engineering teams as a reliability consultant, conducting game days, promoting SRE best practices, and ensuring developers design resilient services from inception.
Qualifications
- Minimum of four years of experience in Site Reliability Engineering, DevOps, or platform engineering roles with responsibility for production environments at scale.
- Practical expertise in Kubernetes operations including deployment, scaling, networking, and troubleshooting in a live context.
- Proficient with infrastructure as code, especially Terraform or Pulumi, focused on a major cloud platform like AWS.
- Experience with observability tools such as Prometheus and Grafana, proficient in creating effective dashboards and alerting systems.
- Competency in at least one programming or scripting language such as Python, Go, or Bash to automate and develop tooling.
- Possess an AWS certification such as Solutions Architect, DevOps Engineer, or SysOps Administrator.
Benefits and Perks
- Competitive salary aligned with your experience and skillset, including bonuses based on performance.
- Comprehensive benefits package, encompassing accommodation, meal allowances, and support with the work visa process.
- Work-life balance supported by paid holidays and bonuses during New Year celebrations.
- Provision of advanced equipment, including a MacBook and an iPhone to maintain productivity.
- Engagement in an energetic and inclusive company culture that encourages personal and professional growth.
Tools & software
How they work
Teamwork & Collaboration
Problem Solving
Accountability
Resilience