Expert Site Reliability Engineer
Riyadh, Riyadh Province, Saudi Arabia · Full Time
Be the first to apply
- Experience
- 5+ yrs
- Salary
- —
- Openings
- 1
- Posted
- 5 days ago
- Work mode
- In office
- Education
- Bachelor's degree
- Resume
- Required to apply
Where you'll work
Sign in to tell us what does and doesn't work for you here — it sharpens every match we show you.
Job description
Role Overview
The position aims to enhance the reliability, availability, scalability, and operational resilience of key technology services through advanced software engineering, automation, observability, and reliability engineering practices.
Key Responsibilities
- Implement and champion sophisticated reliability engineering methodologies for critical technological solutions.
- Define, establish, and monitor service-level indicators (SLIs), objectives (SLOs), and set reliability benchmarks.
- Create automation workflows to minimize manual operational tasks and bolster system resilience.
- Build and refine monitoring, observability, alerting, and incident detection frameworks.
- Lead complex production incident analyses and technical resolution efforts.
- Conduct root cause investigations and enforce permanent remedial and preventive strategies.
- Architect solutions aimed at enhancing availability, scalability, capacity, and disaster resilience.
- Identify potential reliability vulnerabilities and suggest architectural and engineering improvements.
- Oversee performance engineering and capacity strategizing for vital services.
- Provide expert technical mentorship and guidance on site reliability engineering best practices.
- Advocate for automation and engineering approaches that minimize operational toil and boost overall service reliability.
Required Qualifications and Experience
- Bachelor’s degree in Computer Science, Software Engineering, IT, or related disciplines.
- Minimum of five years’ experience in Site Reliability Engineering, DevOps, Platform Engineering, or comparable roles.
- Proficient with cloud platforms, Kubernetes, and managing production environments.
- Deep understanding of monitoring systems, observability, alerting mechanisms, SLIs, SLOs, and related reliability metrics.
- Hands-on experience with automation, scripting, continuous integration/continuous deployment (CI/CD), and Infrastructure as Code (e.g., Terraform).
- Demonstrated expertise in handling complex incidents, performing troubleshooting, and conducting root cause analyses.
- Strong grasp of principles related to high availability, scalability, performance engineering, capacity planning, and disaster recovery protocols.
- Track record of driving reliability improvements and reducing operational burdens via automation techniques.
- Excellent analytical thinking, problem-solving capabilities, and leadership skills within a technical context.
- Experience in the Banking, FinTech, or Payment sectors is advantageous.
Minimum education
Bachelor's Degree
Skills
Tools & software
Kubernetes
required
CI/CD
required
How they work
Problem Solving
Leadership