TAWANTECH

Expert Site Reliability Engineer

TAWANTECH

Riyadh, Riyadh Province, Saudi Arabia · Full Time

Be the first to apply

Experience
5+ yrs
Salary
Openings
1
Posted
5 days ago
Work mode
In office
Education
Bachelor's degree
Resume
Required to apply

Where you'll work

Sign in to tell us what does and doesn't work for you here — it sharpens every match we show you.

Job description

Role Overview

The position aims to enhance the reliability, availability, scalability, and operational resilience of key technology services through advanced software engineering, automation, observability, and reliability engineering practices.

Key Responsibilities

  • Implement and champion sophisticated reliability engineering methodologies for critical technological solutions.
  • Define, establish, and monitor service-level indicators (SLIs), objectives (SLOs), and set reliability benchmarks.
  • Create automation workflows to minimize manual operational tasks and bolster system resilience.
  • Build and refine monitoring, observability, alerting, and incident detection frameworks.
  • Lead complex production incident analyses and technical resolution efforts.
  • Conduct root cause investigations and enforce permanent remedial and preventive strategies.
  • Architect solutions aimed at enhancing availability, scalability, capacity, and disaster resilience.
  • Identify potential reliability vulnerabilities and suggest architectural and engineering improvements.
  • Oversee performance engineering and capacity strategizing for vital services.
  • Provide expert technical mentorship and guidance on site reliability engineering best practices.
  • Advocate for automation and engineering approaches that minimize operational toil and boost overall service reliability.

Required Qualifications and Experience

  • Bachelor’s degree in Computer Science, Software Engineering, IT, or related disciplines.
  • Minimum of five years’ experience in Site Reliability Engineering, DevOps, Platform Engineering, or comparable roles.
  • Proficient with cloud platforms, Kubernetes, and managing production environments.
  • Deep understanding of monitoring systems, observability, alerting mechanisms, SLIs, SLOs, and related reliability metrics.
  • Hands-on experience with automation, scripting, continuous integration/continuous deployment (CI/CD), and Infrastructure as Code (e.g., Terraform).
  • Demonstrated expertise in handling complex incidents, performing troubleshooting, and conducting root cause analyses.
  • Strong grasp of principles related to high availability, scalability, performance engineering, capacity planning, and disaster recovery protocols.
  • Track record of driving reliability improvements and reducing operational burdens via automation techniques.
  • Excellent analytical thinking, problem-solving capabilities, and leadership skills within a technical context.
  • Experience in the Banking, FinTech, or Payment sectors is advantageous.

Minimum education

Bachelor's Degree

Tools & software

Kubernetes required CI/CD required

How they work

Problem Solving Leadership
🤖
Online · instant AI help
Broxer