SmartChoice International GCC

Senior DevOps Engineer

SmartChoice International GCC

Riyadh, Riyadh Province, Saudi Arabia · Full Time

Be the first to apply

Experience
8+ yrs
Salary
Openings
1
Posted
12 seconds ago
Work mode
In office
Resume
Required to apply

Where you'll work

Sign in to tell us what does and doesn't work for you here — it sharpens every match we show you.

Job description

About the Role

We are collaborating with a thriving technology company to recruit a Senior DevOps Engineer responsible for developing and managing the infrastructure supporting a rapidly expanding tech platform. This engaging position covers cloud infrastructure, automation, CI/CD processes, container orchestration, security, monitoring, and maintaining production stability.

Key Responsibilities

  • Architect and sustain scalable, reproducible infrastructures through infrastructure as code.
  • Oversee CI/CD pipelines encompassing build, testing, security assessments, and deployment to production.
  • Administer containerized applications running on Kubernetes and cloud compute platforms.
  • Design secure deployment methods incorporating rollout, rollback, and recovery protocols.
  • Handle cloud networking elements such as API gateways, load balancers, DNS setup, certificates, and inter-service connectivity.
  • Implement identity management, secrets handling, and least-privileged access strategies.
  • Operate production-grade databases ensuring high availability, replication, backups, and recovery capabilities.
  • Set up monitoring systems, alerts, and key operational metrics for infrastructure and application layers.
  • Lead incident management efforts and drive continuous enhancements to reliability, security, and operational resilience.
  • Optimize resource usage and related costs, especially concerning compute-intensive workloads.
  • Contribute to building reusable platform services empowering engineering teams to independently deploy and manage their applications.

Required Qualifications and Experience

  • Minimum of 8 years in DevOps, Site Reliability Engineering, infrastructure, or platform engineering roles.
  • Extensive hands-on expertise with AWS services including compute, networking, identity and access management (IAM), and storage.
  • Proficiency with infrastructure-as-code tools like Terraform or OpenTofu, including developing reusable modules and managing state.
  • Experience utilizing configuration management tools such as Ansible or similar.
  • Strong background in CI/CD pipeline tools, especially GitHub Actions or comparable platforms.
  • Skilled in containerization technologies including Docker, Kubernetes, Helm; familiarity with Amazon EKS preferred.
  • Knowledge of GitOps practices involving ArgoCD or Flux deployment tools.
  • Experience with API gateways managing authentication, routing, and traffic rate limiting.
  • Hands-on production experience with PostgreSQL, AWS RDS/Aurora, and Redis or equivalents.
  • Familiarity with secrets management solutions such as Vault, AWS Secrets Manager, or Parameter Store.
  • Competence in observability platforms like Prometheus, Grafana, AWS CloudWatch, and OpenTelemetry or their equivalents.
  • Comfort working within Linux environments using Bash scripting and Python programming.
  • Strong grasp of security best practices, system reliability, incident handling, and production operations.

Preferred Advanced Infrastructure Experience

Experience with AI, machine learning, or high-demand compute workloads in a production setting is highly valued. Relevant expertise includes:

  • Managing GPU-based infrastructure and scheduling workloads efficiently.
  • Operating model-serving or inference platforms.
  • Implementing autoscaling and capacity planning for compute-heavy applications.
  • Monitoring performance, utilization, and cost management of high-performance environments.
  • Familiarity with technologies such as vLLM, Triton, TGI, KServe, Ray Serve, Amazon Bedrock, or SageMaker.

Technical Stack

Key technologies employed include AWS, Terraform/OpenTofu, Ansible, GitHub Actions, Docker, Kubernetes/EKS, Helm, ArgoCD/Flux, API gateways, PostgreSQL, RDS/Aurora, Redis, Vault/Secrets Manager, Prometheus, Grafana, CloudWatch, OpenTelemetry, Linux, Bash, and Python.

Candidate Profile

We seek an engineer with profound technical expertise and sound operational judgment capable of calmly addressing complex production issues. The ideal candidate should understand the ramifications of infrastructure modifications and focus on designing systems prioritized for dependability, security, and cost-efficiency.

Level

Senior

Tools & software

Docker required Kubernetes required PostgreSQL required Redis required Prometheus required Amazon Web Services AWS required Ansible required GitHub Actions required AWS RDS required Terraform required
🤖
Online · instant AI help
Broxer