- Experience
- 6–10 yrs
- Salary
- —
- Openings
- 1
- Posted
- 5 days ago
- Work mode
- Work from home
- Eligibility
- Applicants must have authorization to work in their home country; visa sponsorship is not provided.
- Resume
- Required to apply
Sign in to tell us what does and doesn't work for you here — it sharpens every match we show you.
Job description
Overview
This position represents a senior role within a globally distributed infrastructure team, supporting critical production systems for a high-throughput platform based in the United Arab Emirates. The role involves comprehensive ownership of system reliability and operational performance, focusing on cloud infrastructure, observability, incident response, performance management, capacity planning, and safe deployment practices.
Key Responsibilities
- Take full responsibility for the reliability of essential production systems end-to-end, covering instrumentation, setting reliability objectives, daily operations, and ensuring performance under real traffic conditions.
- Establish and uphold service level indicators (SLIs), service level objectives (SLOs), and error budgets, linking reliability metrics to engineering priorities and resource allocation.
- Enhance alert systems by increasing actionable alerts, minimizing false alarms, and identifying gaps in the detection of customer-impacting issues.
- Lead incident response efforts for severe incidents, conducting thorough investigations across services, restoring functionality, and facilitating actionable postmortems.
- Design and maintain infrastructure that is secure, fault-tolerant, cost-effective, and resilient to failure modes such as timeouts, retries, backpressure, and graceful degradation.
- Conduct capacity planning and performance analysis through load tests, profiling, saturation studies, and proactive headroom management.
- Improve deployment safety with strategies like progressive delivery, automated rollbacks, validation steps before production, and robust deployment processes.
- Manage infrastructure as code primarily using Terraform, ensuring consistency with the overall service architecture.
- Develop production-grade software and operational tooling that reduce manual efforts and streamline service ownership.
- Organize failure testing exercises such as chaos engineering and game days to uncover system vulnerabilities and define safe operating boundaries.
- Collaborate with product engineering teams to guarantee production readiness, including capacity evaluation, failure handling, rollback plans, runbook creation, and managing on-call transitions.
- Participate actively in on-call rotations, continually enhancing operational procedures to improve runbook clarity, escalation mechanisms, and reduce pager fatigue.
- Apply a security-first mindset to all engineering practices by identifying vulnerabilities in implementations and during code reviews.
- Mentor fellow engineers through code reviews, paired programming sessions, and feedback on design decisions.
Required Qualifications and Experience
- Between 6 to 10 years of experience in site reliability engineering, production engineering, infrastructure, or backend development, preferably with substantial cloud environment exposure and accountability for production systems.
- Demonstrated ability to fully own system lifecycle from design and implementation through live operations and continuous enhancements.
- Expertise in formulating, deploying, and maintaining SLIs, SLOs, and error budgets.
- Proven experience in managing high-impact, customer-facing production incidents and refining incident management protocols at an organizational level.
- In-depth knowledge of distributed system failure patterns within high-throughput, low-latency settings, including cache/database saturation, cascading failures, retry storms, capacity bottlenecks, performance degradation, and load shedding.
- Solid understanding of cloud infrastructure foundations, including networking, container orchestration (Kubernetes/EKS), load balancing, and distributed system design.
- Comprehensive experience managing infrastructure using Infrastructure as Code tools like Terraform.
- Proficiency in at least one programming language such as Go or Python, with proven ability to build and deploy production software.
- Hands-on experience with observability stacks like Datadog, Prometheus, Grafana, or OpenTelemetry, including instrumentation of systems for monitoring.
- Operational experience with Redis or ElastiCache systems, including cluster management, failover handling, eviction policies, and scaling methodologies.
- Strong software development practices including version control, peer review, thorough testing, and safe deployment techniques.
- High degree of autonomy with solid ownership skills, able to progress with minimal defined requirements.
- Balanced, pragmatic approach to reliability focusing on delivering value while managing engineering risks thoughtfully.
- Excellent command of English for technical documentation, reviews, incident communication, and postmortem analysis.
- Experience using AI tools in engineering workflows for tasks like incident investigation and telemetry analysis, with sound judgment regarding their capabilities.
Benefits and Additional Information
- Competitive salary aligned with industry standards and candidate experience.
- Fully remote work arrangement.
- Collaborate with a globally distributed engineering team.
- Opportunity to take significant responsibility for the reliability and operational excellence of production-grade infrastructure.
- Work on high-throughput and low-latency systems with direct impact on customer experience.
- Gain exposure to modern cloud platforms, observability tools, infrastructure as code, and AI-enhanced workflows.
- Influence engineering standards and production readiness across multiple product teams.
- Work within a culture that values ownership, continuous improvement, collaboration, and knowledge sharing.
- Inclusive and diverse workplace culture.
- Applicants must be authorized to work in their home country without visa sponsorship.
Additional Details
This opportunity is managed by a partner company who oversees all applications and subsequent steps. The recruitment process includes an AI-powered candidate matching system designed to provide objective and fair evaluation, ensuring only top candidates are shortlisted for the company to make final hiring decisions. Candidates’ personal data will be processed in compliance with relevant data protection laws, including GDPR, and AI is used only as a support tool in the recruitment process without replacing human judgment.
Level
Senior