Site Reliability Engineer
Abu Dhabi Emirate, United Arab Emirates · Full Time
Be the first to apply
- Experience
- 4–7 yrs
- Salary
- —
- Openings
- 1
- Posted
- 1 week ago
- Work mode
- In office
- Education
- Bachelor's degree
- Resume
- Required to apply
Sign in to tell us what does and doesn't work for you here — it sharpens every match we show you.
Job description
Company Overview
Open Innovation AI is a global tech firm focused on creating cutting-edge tools to manage AI workloads efficiently. Its main product, the Open Innovation Cluster Manager (OICM), streamlines complex AI operations across various hardware, including multiple GPU models and accelerators. This platform is designed to be hardware-independent, enabling easy scalability and integration for enterprise-level AI solutions. The company aims to simplify AI workload management, reduce operational costs, speed up time to market, and enhance returns on AI investments for organizations of all sizes.
Role Overview
The Site Reliability Engineer plays a crucial role in supporting and sustaining Open Innovation AI products deployed in customer environments, especially secure and air-gapped on-premises setups. The role demands strong diagnostic capabilities across hardware, Linux operating systems, Kubernetes, middleware, and application layers. The engineer will manage incidents, troubleshoot complex issues, and maintain high service availability by applying deep knowledge of the products and operational procedures including Incident, Change, and Problem Management. Understanding the product architecture and real-world customer application is essential.
Key Responsibilities
- Deploy and support customer production systems, focusing on environments with restricted network connectivity or isolated infrastructure.
- Maintain system availability, perform upgrades, and ensure the stability of AI platform solutions end-to-end.
- Analyze and fix technical incidents spanning hardware, Linux OS, Kubernetes clusters, container-based services, middleware, and platform elements.
- Investigate logs, system behaviors, and application outputs to identify root causes and restore normal operations.
- Implement authorized changes such as system updates and configuration modifications following formal Change Management protocols.
- Maintain thorough knowledge of OICM and other Open Innovation product architectures, service dependencies, and customer usage patterns.
- Collaborate with first-level support teams to provide technical insights, clarify issues, and facilitate accurate ticket handling.
- Escalate advanced, code-level, or defect-related issues to senior engineering teams with detailed diagnostics and methodical analyses.
- Perform on-site platform health checks including Kubernetes status, service integrity, resource utilization, and readiness assessments.
- Work jointly with Systems Engineering to diagnose and resolve performance bottlenecks across compute, storage, networking, and Kubernetes infrastructure, integrating solutions into the product and operational workflows.
- Create and continually update technical documents, including standard operating procedures, runbooks, troubleshooting guides, and known issue repositories.
- Participate actively in post-incident reviews by providing technical expertise and recommending preventive measures.
- Ensure compliance with Incident, Change, and Problem Management standards and processes.
Required Qualifications and Experience
- Bachelor’s degree in Computer Science, Information Technology, Engineering, or related discipline.
- Between 4 and 7 years of professional experience in Site Reliability Engineering, DevOps, Infrastructure Operations, or Platform Engineering particularly in on-premises or secure environments.
- Advanced Linux system administration skills including troubleshooting, log analysis, managing services, and performance tuning.
- Practical experience working with Kubernetes, container runtimes, and distributed systems in on-prem installations.
- Comprehensive understanding of enterprise compute, storage, networking, and virtualization layers.
- Hands-on experience with middleware and data-layer components like Kafka, Redis, PostgreSQL, or analogous technologies in distributed on-prem environments.
- Well-versed in ITIL-aligned operational frameworks with experience in structured management processes.
- Proven ability to diagnose complex multi-layered technical issues effectively.
- Experience in secure, restricted, or air-gapped environments considered a significant advantage.
- Strong analytical thinking, effective communication skills, and a methodical approach to troubleshooting challenges.
- Capability to develop clear and comprehensive technical documentation including SOPs, runbooks, and investigative reports.
- Certifications such as Red Hat Certified System Administrator/Engineer (RHCSA/RHCE), Certified Kubernetes Administrator/Developer/Security Specialist (CKA/CKAD/CKS) are preferred.
Minimum education
Bachelor's Degree