- Experience
- 10+ yrs
- Salary
- USD 177,000 – USD 240,000 / year
- Openings
- 1
- Posted
- 1 week ago
- Work mode
- Work from home
- Resume
- Required to apply
Sign in to tell us what does and doesn't work for you here — it sharpens every match we show you.
Job description
About the Role
This opportunity is to serve as the inaugural dedicated Site Reliability Engineer (SRE) for a partner organization based in Saudi Arabia, functioning within a fully remote, global engineering environment. The role focuses on leading and establishing robust reliability engineering standards across multiple engineering teams and key production systems.
Key Responsibilities
- Define and implement Service Level Indicators (SLIs) and Service Level Objectives (SLOs) for essential production request pathways, ensuring that reliability targets are measurable, transparent, and integrated with engineering decisions.
- Promote and apply error budgets to judiciously balance reliability investments against product and feature delivery.
- Develop and sustain reliability metrics that enable engineering leadership to monitor progress and allocate resources aptly.
- Enhance the incident management process, encompassing detection, response, communication, escalation, postmortems, and follow-up actions.
- Collaborate with infrastructure teams to improve alert accuracy, anomaly detection, escalation protocols, and operational tooling.
- Conduct reliability evaluations for high-risk deployments and new services, focusing on readiness, capacity planning, failure modes, rollback approaches, and operational risk mitigation.
- Introduce and lead failure testing initiatives such as game days and chaos engineering exercises to proactively uncover vulnerabilities and confirm safe operational thresholds.
- Engage with engineering teams to address complex reliability challenges, fostering improved practices and clear ownership.
- Mentor staff and lead engineers to cultivate reliability champions and promote decentralized SRE ownership.
- Create streamlined operational standards covering production readiness, on-call routines, runbooks, change safety, and service operability.
- Partner with architects and technical leads to embed reliability and fault tolerance into system design phases rather than post-deployment remediation.
- Maintain hands-on involvement during live incidents and investigations, developing tools, dashboards, automation, and exemplar implementations where appropriate.
- Advance the use of artificial intelligence for incident investigation, telemetry analysis, postmortem development, runbook creation, observability, and reliability tooling.
- Organize operational data, alerts, dashboards, and runbooks to facilitate safe interpretation and actions by both engineers and AI agents.
- Contribute direct code and infrastructure modifications to production environments supporting reliability improvements.
Candidate Requirements
- Minimum of 10 years in engineering roles, including at least 3 years specializing in SRE, production engineering, or a reliability-oriented staff engineer capacity impacting multiple teams.
- Proven track record for managing platform-level or organizational reliability rather than individual service ownership.
- Expertise in designing and operationalizing SLIs, SLOs, and error budgets, with successful organizational adoption.
- Substantial experience in incident leadership, including coordination of high-impact, customer-facing incidents and conducting effective postmortems with measurable outcomes.
- Thorough understanding of distributed system failure modes such as cache/database saturation, cascading failures, retry storms, capacity limits, graceful degradation, and load shedding.
- Strong practical knowledge of Kubernetes, AWS environments, and modern observability platforms (e.g., Datadog or similar).
- Proficient in writing and reviewing production code in Go, TypeScript, or comparable languages, alongside infrastructure as code practices.
- Ability to influence engineering teams cross-functionally without formal authority and drive cultural change towards reliability.
- Experienced in coaching and mentoring engineers to build reliability ownership competencies.
- Excellent verbal and written communication abilities to articulate technical trade-offs, risks, incidents, and priorities to diverse audiences including executives.
- Preference for asynchronous workflows with clearly documented decisions and communications.
- Hands-on familiarity with AI applications in incident investigation, telemetry insight, runbook/postmortem development, and engineering tooling.
- Insight into structuring operational assets like alerts, dashboards, and runbooks to support AI-assisted diagnostics and interventions.
- Balanced, pragmatic approach weighing operational risk against engineering investment, delivery timelines, and overarching business goals.
- Additional assets include experience in fraud detection, identity management, payments, adversarial system environments, multi-region or cell-based architectures, and failure isolation methods.
- Experience managing large scale Elasticsearch, Redis, DynamoDB, or Kafka deployments and understanding corresponding failure mechanisms is advantageous.
- Awareness of FinOps principles and cloud infrastructure cost versus reliability trade-offs adds value.
- Must be eligible to work legally in Saudi Arabia; no visa sponsorship provided.
Benefits and Work Environment
- Fully remote position with flexible work location.
- Opportunity to be a pioneer in establishing SRE frameworks organization-wide.
- High autonomy with significant authority to define engineering reliability standards and operational procedures.
- Collaborative cross-team and cross-functional exposure including leadership, architecture, and infrastructure domains.
- Chance to shape AI-driven reliability and operational tooling practices for the future.
- Strong focus on professional development, leadership growth, coaching, and knowledge sharing.
- Inclusive global remote culture that embraces diverse perspectives and backgrounds.
- For employees based in the United States, the compensation range is approximately $177,000 to $240,000 USD, varying based on skills, experience, education, certifications, and location. Salaries for other regions may differ.
- Remote work participation is subject to applicable regulatory and security requirements based on the candidate's jurisdiction.
Additional Information
This role is recruited on behalf of a partner organization that handles all application processing and subsequent steps. The partner uses AI-supported candidate matching to ensure fair and objective evaluation against core requirements, but final hiring decisions rest solely with their internal team.
Level
Mid