Jobgether

Staff Site Reliability Engineer

Jobgether

Remote · Full Time

Be the first to apply

Experience
10+ yrs
Salary
USD 177,000 – USD 240,000 / year
Openings
1
Posted
1 week ago
Work mode
Work from home
Resume
Required to apply

Sign in to tell us what does and doesn't work for you here — it sharpens every match we show you.

Job description

About the Role

We are seeking a Staff Site Reliability Engineer to join a globally distributed remote engineering group based in the United Arab Emirates. This pivotal role focuses on driving reliability leadership across multiple engineering teams and critical production systems. As the first dedicated SRE on the team, you will establish and embed reliability practices throughout the organization, working closely with engineering leaders, architects, and infrastructure experts.

Key Responsibilities

  • Define and implement Service Level Indicators (SLIs) and Service Level Objectives (SLOs) for key production workflows to ensure reliability targets are visible and actionable within the engineering organization.
  • Advocate for and manage error budgets that balance reliability investments alongside product development and feature rollouts.
  • Develop and maintain reliability metrics that guide engineering leadership decisions and highlight areas needing improvement.
  • Enhance the entire incident management process including detection, communication, escalation, response, and conducting postmortems with follow-up actions.
  • Collaborate with infrastructure teams to improve alert accuracy, anomaly detection, escalation protocols, and shared operational tooling.
  • Lead reliability reviews for high-impact changes and new system deployments, including production readiness, capacity planning, failure scenarios, rollback procedures, and risk assessments.
  • Introduce and facilitate deliberate failure testing, game days, and chaos engineering exercises to reveal system vulnerabilities and validate operational limits before problems arise.
  • Partner directly with engineering teams to tackle complex reliability challenges, instilling durable practices and clear ownership frameworks.
  • Coach staff and lead engineers to foster reliability advocacy within teams and promote distributed SRE responsibilities.
  • Create efficient and repeatable operational standards addressing production readiness, on-call best practices, runbooks, change safety, and service operability.
  • Work alongside architects and technical leads to embed reliability and failure resilience in system designs from the outset rather than as post-deployment fixes.
  • Remain hands-on during production incidents, contributing by building automation, dashboards, tooling, and reference implementations as necessary.
  • Promote the use of artificial intelligence to support incident investigation, telemetry analysis, postmortem compilation, runbook creation, observability enhancement, and reliability tooling innovations.
  • Structure operational data, alerts, dashboards, and runbooks to be understandable and actionable by both engineers and AI agents safely.
  • Take active part in producing fixes and improvements through code and infrastructure updates, not limited to advisory roles.

Candidate Requirements

  • Over 10 years of engineering experience including at least 3 years in site reliability engineering, production engineering, or a related reliability-focused Staff Engineer role spanning multiple teams.
  • Proven track record of managing reliability for platforms or entire organizations beyond individual services.
  • In-depth experience designing and applying SLIs, SLOs, and error budgets, with a successful record of driving organization-wide adoption.
  • Extensive incident management proficiency, including leadership during severe customer-impacting incidents and conducting meaningful postmortems that lead to tangible improvements.
  • Advanced understanding of failure modes in distributed systems, such as database/cache saturation, cascading failures, retry storms, capacity limitations, graceful degradation, and load shedding.
  • Strong practical expertise with Kubernetes, AWS, and advanced observability platforms like Datadog or equivalents.
  • Ability to write and maintain production code in Go, TypeScript, or similar languages alongside infrastructure-as-code proficiency.
  • Demonstrated success influencing engineering practices and culture changes without formal authority.
  • Excellent coaching and mentoring capabilities, with demonstrated development of engineers into reliability champions.
  • Exceptional communication skills, capable of conveying incident reports, risks, trade-offs, and priorities clearly across technical and executive audiences.
  • Preference for asynchronous, well-documented decision-making, and transparent technical communication.
  • Experience leveraging AI tools for incident investigation, telemetry evaluation, runbook and postmortem documentation, and automation tooling.
  • Knowledge of operational data structuring to support safe AI-driven diagnosis and interventions.
  • Balanced approach to reliability, considering operational risk, engineering costs, delivery timelines, and business objectives.
  • Bonus qualifications include experience in fraud detection, identity management, payments, or real-time adversarial systems; familiarity with multi-region or failure-isolation architectures; operational knowledge of Elasticsearch, Redis, DynamoDB, Kafka at scale; and understanding of FinOps and cloud reliability-cost trade-offs.
  • Applicants must have authorization to work in the United Arab Emirates; visa sponsorship is not provided.

Benefits and Work Environment

  • Fully remote position with flexibility for global collaboration.
  • Opportunity to pioneer reliability practices as the organization’s first dedicated SRE.
  • High autonomy to define standards, coaching programs, and scalable reliability frameworks.
  • Broad influence across multiple teams and critical infrastructure.
  • Close interaction with engineering leadership, architects, infrastructure, and product teams.
  • Engage in shaping AI-assisted reliability methods and future production operations.
  • Support for professional development and leadership growth in technical domains.
  • Inclusive, diverse, and globally dispersed engineering culture.
  • US employees have a compensation range between $177,000 and $240,000, with variation by experience and market; other regions may differ.
  • Remote work eligibility depends on regulatory and security requirements at candidate location.

Additional Information

This role is listed on behalf of a partner organization that handles the application process and next steps. The hiring company manages interviews and decisions directly.

Level

Mid

Tools & software

AWS Amazon Web Services AWS required

How they work

Communication Leadership Persuasion & Influence
🤖
Online · instant AI help
Broxer