Jobgether

Staff Site Reliability Engineer

Jobgether

Remote · Full Time

Be the first to apply

Experience
10+ yrs
Salary
USD 177,000 – USD 240,000 / year
Openings
1
Posted
1 week ago
Work mode
Work from home
Resume
Required to apply

Sign in to tell us what does and doesn't work for you here — it sharpens every match we show you.

Job description

About the Role

This opportunity is to serve as the inaugural dedicated Site Reliability Engineer (SRE) for a partner organization based in Saudi Arabia, functioning within a fully remote, global engineering environment. The role focuses on leading and establishing robust reliability engineering standards across multiple engineering teams and key production systems.

Key Responsibilities

  • Define and implement Service Level Indicators (SLIs) and Service Level Objectives (SLOs) for essential production request pathways, ensuring that reliability targets are measurable, transparent, and integrated with engineering decisions.
  • Promote and apply error budgets to judiciously balance reliability investments against product and feature delivery.
  • Develop and sustain reliability metrics that enable engineering leadership to monitor progress and allocate resources aptly.
  • Enhance the incident management process, encompassing detection, response, communication, escalation, postmortems, and follow-up actions.
  • Collaborate with infrastructure teams to improve alert accuracy, anomaly detection, escalation protocols, and operational tooling.
  • Conduct reliability evaluations for high-risk deployments and new services, focusing on readiness, capacity planning, failure modes, rollback approaches, and operational risk mitigation.
  • Introduce and lead failure testing initiatives such as game days and chaos engineering exercises to proactively uncover vulnerabilities and confirm safe operational thresholds.
  • Engage with engineering teams to address complex reliability challenges, fostering improved practices and clear ownership.
  • Mentor staff and lead engineers to cultivate reliability champions and promote decentralized SRE ownership.
  • Create streamlined operational standards covering production readiness, on-call routines, runbooks, change safety, and service operability.
  • Partner with architects and technical leads to embed reliability and fault tolerance into system design phases rather than post-deployment remediation.
  • Maintain hands-on involvement during live incidents and investigations, developing tools, dashboards, automation, and exemplar implementations where appropriate.
  • Advance the use of artificial intelligence for incident investigation, telemetry analysis, postmortem development, runbook creation, observability, and reliability tooling.
  • Organize operational data, alerts, dashboards, and runbooks to facilitate safe interpretation and actions by both engineers and AI agents.
  • Contribute direct code and infrastructure modifications to production environments supporting reliability improvements.

Candidate Requirements

  • Minimum of 10 years in engineering roles, including at least 3 years specializing in SRE, production engineering, or a reliability-oriented staff engineer capacity impacting multiple teams.
  • Proven track record for managing platform-level or organizational reliability rather than individual service ownership.
  • Expertise in designing and operationalizing SLIs, SLOs, and error budgets, with successful organizational adoption.
  • Substantial experience in incident leadership, including coordination of high-impact, customer-facing incidents and conducting effective postmortems with measurable outcomes.
  • Thorough understanding of distributed system failure modes such as cache/database saturation, cascading failures, retry storms, capacity limits, graceful degradation, and load shedding.
  • Strong practical knowledge of Kubernetes, AWS environments, and modern observability platforms (e.g., Datadog or similar).
  • Proficient in writing and reviewing production code in Go, TypeScript, or comparable languages, alongside infrastructure as code practices.
  • Ability to influence engineering teams cross-functionally without formal authority and drive cultural change towards reliability.
  • Experienced in coaching and mentoring engineers to build reliability ownership competencies.
  • Excellent verbal and written communication abilities to articulate technical trade-offs, risks, incidents, and priorities to diverse audiences including executives.
  • Preference for asynchronous workflows with clearly documented decisions and communications.
  • Hands-on familiarity with AI applications in incident investigation, telemetry insight, runbook/postmortem development, and engineering tooling.
  • Insight into structuring operational assets like alerts, dashboards, and runbooks to support AI-assisted diagnostics and interventions.
  • Balanced, pragmatic approach weighing operational risk against engineering investment, delivery timelines, and overarching business goals.
  • Additional assets include experience in fraud detection, identity management, payments, adversarial system environments, multi-region or cell-based architectures, and failure isolation methods.
  • Experience managing large scale Elasticsearch, Redis, DynamoDB, or Kafka deployments and understanding corresponding failure mechanisms is advantageous.
  • Awareness of FinOps principles and cloud infrastructure cost versus reliability trade-offs adds value.
  • Must be eligible to work legally in Saudi Arabia; no visa sponsorship provided.

Benefits and Work Environment

  • Fully remote position with flexible work location.
  • Opportunity to be a pioneer in establishing SRE frameworks organization-wide.
  • High autonomy with significant authority to define engineering reliability standards and operational procedures.
  • Collaborative cross-team and cross-functional exposure including leadership, architecture, and infrastructure domains.
  • Chance to shape AI-driven reliability and operational tooling practices for the future.
  • Strong focus on professional development, leadership growth, coaching, and knowledge sharing.
  • Inclusive global remote culture that embraces diverse perspectives and backgrounds.
  • For employees based in the United States, the compensation range is approximately $177,000 to $240,000 USD, varying based on skills, experience, education, certifications, and location. Salaries for other regions may differ.
  • Remote work participation is subject to applicable regulatory and security requirements based on the candidate's jurisdiction.

Additional Information

This role is recruited on behalf of a partner organization that handles all application processing and subsequent steps. The partner uses AI-supported candidate matching to ensure fair and objective evaluation against core requirements, but final hiring decisions rest solely with their internal team.

Level

Mid

Tools & software

Kubernetes · 2 to 5 years required Amazon Web Services AWS required

How they work

Communication Leadership Decision Making

Leave it if you'd like a reply — we won't use it for anything else.

Click to browse, drag & drop, or paste a screenshot

PNG, JPG, GIF, MP4, WebM, MOV · Max 20MB each · Up to 5 files

🤖
Online · instant AI help
Broxer