Senior Site Reliability Engineer
Dubai, United Arab Emirates · Full Time
Be the first to apply
- Experience
- 6+ yrs
- Salary
- —
- Openings
- 1
- Posted
- 3 weeks ago
- Work mode
- In office
- Resume
- Required to apply
Where you'll work
Sign in to tell us what does and doesn't work for you here — it sharpens every match we show you.
Job description
About the Role
AW Connect is partnering with a Dubai-based digital asset company aiming to launch a fully regulated platform. We seek a highly skilled Senior Site Reliability Engineer to take charge of the production infrastructure, reliability, incident management, and overall technology operations within a rigorously regulated environment. This role requires a hands-on engineer who actively manages infrastructure rather than purely performing managerial duties.
Key Responsibilities
- Take full ownership of production cloud infrastructure, including Kubernetes clusters, Infrastructure as Code (IaC) practices, and continuous integration/deployment (CI/CD) pipelines.
- Lead incident response efforts, perform root cause analysis (RCA), and oversee system monitoring and observability.
- Manage backup and restore operations alongside business continuity and disaster recovery planning.
- Administer Identity and Access Management (IAM), Privileged Access Management (PAM), and access control workflows.
- Work collaboratively with outsourced cybersecurity partners including Chief Information Security Officer (CISO), Security Operations Center (SOC), and security service providers.
- Coordinate vulnerability assessment efforts, remediation activities, and responses to penetration test findings.
- Maintain required security and operational controls along with audit documentation.
- Provide support for wallet and custody technology controls where necessary.
Required Qualifications and Experience
- Minimum of 6 years hands-on experience in Site Reliability Engineering, DevOps, DevSecOps, platform engineering, or infrastructure engineering roles.
- At least 2 years experience within regulated or audited sectors such as fintech, banking, brokerage, or exchanges.
- Proficient production-level experience managing AWS or comparable cloud platforms.
- Strong practical experience with Kubernetes and Infrastructure as Code implementations.
- Deep understanding of IAM, PAM, and CI/CD methodologies.
- Demonstrated skill managing live incidents and performing detailed root cause analyses.
- Expertise in disaster recovery commitments including defined Recovery Time Objective (RTO), Recovery Point Objective (RPO), and proven backup/restore procedures.
- Track record of working with outsourced security teams like CISOs and SOCs.
- Familiarity with vulnerability management and applying fixes following penetration testing.
- Comfortable generating and maintaining technical documentation and audit evidence.
Preferred Experience
Knowledge of digital asset technologies such as cryptocurrencies, custody/wallet operations, Fireblocks or BitGo platforms, blockchain infrastructure, or KYT tools is beneficial but not mandatory. Candidates with backgrounds in regulated banking, financial technology, brokerage, or financial infrastructure will also be favorably considered.
Candidate Profile
The ideal candidate is a hands-on practitioner who has taken personal responsibility for production environments and has been the primary responder during system failures. This is not a role suited for those who have recently focused exclusively on team management. Joining a small, senior technology team at this launch phase provides direct interaction with the Group CTO and genuine accountability for your domain.
Application Instructions
Please submit a resume that clearly details your personal ownership, implementation experience, and operational responsibilities.
Level
Senior