Systems Arabia

On-Premise DevOps Engineer

Systems Arabia

Riyadh, Riyadh Province, Saudi Arabia · Full Time

Be the first to apply

Experience
6+ yrs
Salary
Openings
1
Posted
6 days ago
Work mode
In office
Resume
Required to apply

Where you'll work

Sign in to tell us what does and doesn't work for you here — it sharpens every match we show you.

Job description

Role Overview

We are seeking an experienced Senior DevOps / Platform Engineer with more than six years of practical experience managing on-premises, self-managed infrastructure and CI/CD systems. The ideal candidate will have proven capabilities in operating OpenShift, Jenkins/CloudBees, Prometheus, Grafana, ELK, and Linux platforms in production environments. Strong Python backend programming skills along with robust system-design expertise are essential. This role demands an autonomous engineer capable of designing, monitoring, troubleshooting, and optimizing infrastructure while delivering high-availability, scalable, and dependable platforms.

Key Responsibilities

  • Develop and maintain CI/CD pipelines leveraging Jenkins/CloudBees, integrating artifact handling and code quality checks.
  • Architect, administer, and sustain CI/CD tooling components, including Jenkins controllers/agents, artifact management systems, and code quality tools.
  • Oversee Jenkins/CI-CD platform upgrades, plugin management, configuration tuning, backups, high availability settings, and access control policies.
  • Operate OpenShift clusters in production, managing capacity, upgrades, role-based access, networking, ingress points, and optimizing workloads.
  • Lead infrastructure monitoring efforts through Prometheus and Grafana, creating metrics queries, dashboards, service level objectives, and alert tuning.
  • Implement and manage ELK stack for centralized logging, handling log parsing, indexing, and retention policies.
  • Analyze resource consumption to optimize infrastructure utilization and provide enhancement suggestions to relevant stakeholders.
  • Build custom backend services, RESTful APIs, and integrations using Python.
  • Design and maintain primary-disaster recovery (PR-DR) systems for key infrastructure with replication, failover/failback processes, and scheduled DR testing.
  • Administer related databases including installation, backup/restoration, replication, monitoring, and performance tuning.
  • Identify and resolve technical issues spanning CI/CD failures, OpenShift infrastructure, storage, networking, and platform layers.
  • Create and upkeep technical runbooks and detailed platform documentation.
  • Provide mentorship to engineering team members and enhance platform engineering methodologies.

Mandatory Requirements

  • Over six years of experience in DevOps or Platform Engineering roles.
  • Proven operational expertise administering OpenShift clusters in production.
  • Extensive experience with Jenkins/CloudBees including crafting declarative pipelines and Groovy shared libraries.
  • Comprehensive understanding of CI/CD platform frameworks, addressing high availability, distributed build agent management, and scalability aspects.
  • Practical knowledge managing primary and disaster recovery processes involving replication and failover strategies aligned with recovery time and point objectives.
  • Database administration skills with PostgreSQL, MySQL, or similar systems covering backups, restores, and replication.
  • Hands-on experience with artifact repositories and incorporating code quality gates into CI/CD pipelines.
  • Advanced proficiency with Prometheus, Grafana dashboards and alerting, and ELK logging stack.
  • Strong backend Python development background, developing APIs and service integrations.
  • Solid system design acumen, ensuring scalability, reliability, and fault tolerance.
  • Capable Linux system administration and scripting expertise.
  • Familiarity with networking concepts, TLS, secrets management, and security enhancement procedures.
  • Exposure to Infrastructure as Code and GitOps methodologies and tools such as Ansible, Helm, or Argo CD.
  • Excellent analytical and troubleshooting capabilities across infrastructure and applications.

How they work

Problem Solving Leadership Independence
🤖
Online · instant AI help
Broxer