Site Reliability Engineer
TestCrew | Quality Engineering & Software Testing
Al Ahsa, Eastern, Saudi Arabia · Full Time
Be the first to apply
- Experience
- 2–5 yrs
- Salary
- —
- Openings
- 1
- Posted
- 4 days ago
- Work mode
- In office
- Education
- Bachelor's degree
- Resume
- Required to apply
Sign in to tell us what does and doesn't work for you here — it sharpens every match we show you.
Job description
About the Role
TestCrew is looking for Site Reliability Engineers to join a significant enterprise project located in Al Ahsa, Saudi Arabia. The role involves ensuring the availability, performance, reliability, and security of vital IT infrastructure through proactive monitoring, automation, and prompt incident handling. This position demands direct operational involvement working closely with development, operations, and cybersecurity teams to maintain continuous service uptime and improve operational workflows.
Key Responsibilities
- Continuously monitor and oversee enterprise-level infrastructure, applications, and services to guarantee high availability and optimal functioning.
- Operate and maintain monitoring and observability platforms to identify incidents early, analyze trends, and proactively resolve system problems.
- Develop and deploy monitoring approaches, dashboards, alerts, and performance indicators to enhance operational transparency.
- Conduct root cause analyses of incidents and implement preventative solutions to minimize repeat occurrences.
- Automate routine operational tasks, deployments, and maintenance using scripting languages and infrastructure automation frameworks.
- Manage and support CI/CD pipelines and facilitate smooth release management processes.
- Optimize system performance while adhering to service level agreements and industry best practices.
- Assist with business continuity planning, disaster recovery, and building operational resilience.
- Enforce security best practices including system hardening, patch management, and secure workflow procedures.
- Collaborate with cross-functional teams—including development, infrastructure, and security—to tackle complex technical challenges.
- Maintain comprehensive documentation such as runbooks, knowledge bases, and incident history records.
- Participate in major incident handling, conduct post-incident evaluations, and champion continual improvement initiatives.
- Provide on-call support and contribute to shift rotations as needed.
Qualifications
- Bachelor’s degree in Computer Science, IT, Engineering, or a related discipline.
- 2 to 5 years of practical experience in site reliability engineering, DevOps, infrastructure operations, or production support environments.
- Proficient in managing enterprise monitoring and observability tools such as Grafana, Prometheus, Datadog, Instana, or Zabbix.
- Strong troubleshooting and root cause analysis abilities covering infrastructure, applications, and network layers.
- Experience with scripting languages including Python, Bash, and PowerShell.
- Familiarity with infrastructure automation and Infrastructure as Code tools like Ansible and Terraform.
- Knowledge of CI/CD tools such as Jenkins, GitLab CI, or Azure DevOps.
- Understanding of business continuity, disaster recovery, and operational risk principles.
- Awareness of cybersecurity protocols including system hardening, access control, vulnerability management, and patching.
- Knowledge of IT Service Management (ITSM) practices and tools like ServiceNow or Jira Service Management is advantageous.
- Fluent in Arabic and English, with excellent written and verbal communication skills.
Preferred Skills
- ITIL Foundation certification.
- Experience with cloud platforms such as AWS, Azure, or Google Cloud Platform.
- Familiarity with containerization and orchestration (Docker, Kubernetes).
- Exposure to large-scale enterprise or government IT environments.
Personal Attributes
- A proactive mindset focused on ownership and reliability.
- Strong analytical and problem-solving capabilities.
- Ability to maintain composure and systematic approach during critical incidents.
- Excellent communication and teamwork skills.
- Adaptability to work efficiently in fast-paced enterprise settings.
- Availability to work on-site in Al Ahsa and participate in shift/on-call rotations.
Minimum education
Bachelor's Degree