T

Site Reliability Engineer

Thales

Eastern Region · Full Time

Be the first to apply

Experience
4–10 yrs
Salary
—
Openings
1
Posted
3 days ago
Work mode
In office
Resume
Required to apply

Sign in to tell us what does and doesn't work for you here — it sharpens every match we show you.

Job description

About the Role and Company

Based in Khobar, Saudi Arabia, this opportunity is with Thales, a global leader recognized for identity management and data protection solutions essential to digital security. Trusted by over 30,000 organizations worldwide, Thales supports digital trust in sectors like banking, energy, civil aviation, cybersecurity, and defense. In Saudi Arabia, with a history of over 50 years, Thales has established strong local partnerships and aligns closely with Saudi Vision 2030 to promote localisation and build sovereign capabilities through strategic collaborations.

Key Responsibilities

  • Implement and sustain the Sovereign Cloud strategy ensuring compliance with data residency and regulatory obligations across AWS and GCP platforms.
  • Operate and maintain Kubernetes clusters involving upgrades, scaling, and effective workload management.
  • Lead cloud security incident response efforts by investigating, analyzing, and resolving security threats while proactively mitigating vulnerabilities within AWS environments.
  • Enhance and support cloud security posture by embedding security best practices and continuous improvement in AWS cloud strategy.
  • Enable security automation and DevSecOps practices through implementation and tuning of security tools, automating policy responses to ensure secure cloud operations.
  • Design, deploy, and manage cloud infrastructure using Infrastructure as Code with Terraform.
  • Maintain and optimize secure automated CI/CD pipelines leveraging Git and GitLab.
  • Manage secrets and access controls securely using HashiCorp Vault including token lifecycle management and secrets rotation.
  • Troubleshoot complex distributed system issues including infrastructure, networking, containers, and performance bottlenecks.
  • Monitor system observability and reliability leveraging Datadog, define SLIs/SLOs, and drive improvements following SRE principles.
  • Conduct 24x7 incident management, alerting via PagerDuty, root cause analysis, and participate in incident, problem, and change governance.
  • Optimize cloud security and performance through system hardening, vulnerability remediation, tuning, and capacity planning.
  • Drive automation and operational efficiency using Python, Bash, Terraform, and Ansible to reduce manual efforts and boost platform resilience.

Candidate Profile and Experience

  • Minimum 4-5 years of practical experience in DevSecOps and Site Reliability Engineering roles.
  • Proficient in managing Kubernetes clusters with demonstrated hands-on expertise.
  • Experienced in Infrastructure as Code using Terraform and version control/CI-CD pipelines via Git and GitLab.
  • Familiar with HashiCorp Vault, Datadog, PagerDuty, and Confluence tools.
  • Strong foundational knowledge of cloud security principles including identity access management, encryption, container security, network security, and vulnerability handling.
  • Experienced in incident response, change management processes, and root cause analysis.
  • Deep understanding of SRE methodologies including SLIs, SLOs, error budgets, and reliability metrics; capable in scripting with Python and/or Bash to enable automation-first approaches.
  • Solid networking knowledge encompassing TCP/IP, DNS, load balancing, firewalls, VPNs, and private endpoints, with experience or exposure to compliance-controlled environments.
  • Proven ability to lead incident management during critical outages (P1/P2), coordinating cross-disciplinary teams and ensuring timely resolution.
  • Customer and business-oriented approach in managing priorities.
  • Relevant cloud certifications considered an advantage.
  • Mandatory expertise in Linux, Kubernetes, Terraform, Ansible, and Google Cloud Platform.

Additional Information

Thales offers career growth and mobility opportunities globally, fostering flexibility and development across diverse areas of expertise. Joining Thales means becoming part of an 80,000-employee worldwide network where career advancement at home or abroad is encouraged.

Tools & software

Git · 2 to 5 years required Kubernetes · 2 to 5 years required GitLab required Amazon Web Services AWS required Ansible required Google Cloud Platform required Terraform · 2 to 5 years required

How they work

Teamwork & Collaboration Initiative Customer Focus
🤖
Online · instant AI help
Broxer