O

Staff Site Reliability Engineer

Obsidian Security

Manchester, England, United Kingdom · Full Time

Be the first to apply

Experience
5+ yrs
Salary
GBP 124,000 – GBP 141,000 / year
Openings
1
Posted
3 days ago
Work mode
In office
Resume
Required to apply

Where you'll work

Sign in to tell us what does and doesn't work for you here — it sharpens every match we show you.

Job description

About Obsidian Security

Founded in 2017, Obsidian Security was established to address a vital need: protecting SaaS applications where contemporary business activities take place—platforms such as Microsoft 365, Salesforce, and many more. Supported by leading investors like Greylock, Norwest Venture Partners, and IVP, Obsidian has developed a comprehensive SaaS security platform designed to minimize risks, detect and respond to threats, and prevent breaches at their origin. The team comprises experts who have shaped the fields of endpoint and identity security at notable companies such as CrowdStrike, Okta, Cylance, and Carbon Black.

Currently leading the transformation of SaaS security amid the rise of agentic AI, Obsidian is trusted by over 200 organizations worldwide, including global enterprises like Snowflake, T-Mobile, and Pure Storage. Their reach spans North America, Europe, the Middle East, Southeast Asia, Australia, and New Zealand, covering many Fortune 1000 and Global 2000 giants. With expanding global momentum, a growing partner network featuring SentinelOne, Databricks, and Google Cloud, and an upcoming major fundraising event, Obsidian is rapidly scaling towards sustained growth and preparing for an IPO.

Role Overview

As a Staff Site Reliability Engineer at Obsidian, you will shape and lead the overarching reliability vision for a sophisticated, multi-tenant SaaS platform catering to enterprise and financial clientele. Collaborating closely with DevOps and Platform Engineering leadership, you will guide a cohesive reliability strategy that spans the entire organization. Your primary responsibility is to guarantee the platform detects, diagnoses, and communicates system issues proactively to avoid customer impact with consistent reliability.

This role demands hands-on technical expertise in designing and implementing systems that manage complex real-world scenarios, including dependencies on upstream SaaS services, managing infrequent and noisy data signals, and supporting critical enterprise workloads.

Key Duties

  • Develop and spearhead a long-term reliability strategy for services, establishing comprehensive system visibility frameworks and guiding architecture for observability, detection, and fault tolerance.
  • Collaborate across departments to integrate reliability best practices, standardize service level indicators and objectives (SLI/SLOs), and act as a technical expert for escalation.
  • Create intelligent detection solutions such as anomaly detection and connector health monitoring, enabling teams to access self-service observability tools.
  • Establish and refine an incident communication hierarchy, enhance incident response methodologies, and lead postmortem investigations to reinforce system reliability and maintain customer confidence.
  • Engage in hands-on contribution towards system design, monitoring, and debugging within distributed systems and data pipeline infrastructures.

Essential Qualifications

  • Minimum of five years’ experience in site reliability engineering, production engineering, or a closely related discipline.
  • At least three years operating in senior technical leadership roles with responsibilities akin to a staff engineer.
  • Strong proficiency with cloud platforms AWS and/or Google Cloud Platform (GCP).
  • Expertise in Kubernetes orchestration and Helm package management.
  • Experience with observability tools such as Prometheus, Grafana, or equivalents.
  • Familiarity with continuous integration/continuous deployment (CI/CD) systems like GitLab CI/CD or ArgoCD.
  • Proven background in designing and scaling reliability solutions for multi-tenant SaaS environments.
  • Advanced debugging capabilities and system-level understanding spanning distributed microservices and legacy systems.
  • Track record of leading innovation in incident detection, handling, and improving system resilience.
  • A practical, hands-on engineering approach with notable experience building reliability-focused systems rather than mere configuration.

Preferred Experience

  • Background working with B2B SaaS products targeting enterprise or financial industry customers.
  • Knowledge of third-party SaaS connector architectures and data ingestion models.
  • Expertise in developing anomaly detection and smart alerting systems.
  • Experience in designing customer-facing status pages and managing incident communication frameworks.

Why Join Us?

  • Lead and influence company-wide reliability strategies.
  • Architect and develop innovative detection and observability solutions.
  • Address and resolve complex challenges involving distributed systems.
  • Protect mission-critical infrastructure supporting financial sector clients.

Success Metrics

  • System faults identified and resolved before affecting users.
  • Improvement and measurable progress of reliability metrics.
  • Teams utilizing scalable self-service observability tools.
  • Transparent, proactive incident communication that fosters trust.
  • Establishing reliability as a key competitive business advantage.

Employee Benefits

Our compensation packages are structured to support employee well-being professionally and personally. Benefits include competitive pay with equity options, 401k plans, comprehensive healthcare coverage including dental and vision, flexible paid leave, and generous parental leave policies. Additionally, personal and career development resources are offered. Pay ranges are guidelines subject to factors such as location, experience, and skill set. This position also offers potential equity awards and eligibility for incentives based on role.

Diversity & Inclusion

Obsidian is committed to equal opportunity employment, valuing diversity, and selecting candidates based on talent, passion, and empathy. Candidates must provide valid proof of identity and legal work authorization. Accommodations are available upon request. Application data is handled following Obsidian’s confidentiality policies.

Level

Mid

Tools & software

Kubernetes · 2 to 5 years required Prometheus · 2 to 5 years required Amazon Web Services AWS required Grafana Labs Grafana Cloud required Argocd · 2 to 5 years required Google Cloud Platform · 2 to 5 years required

How they work

Teamwork & Collaboration Problem Solving Leadership Initiative Strategic Thinking

Leave it if you'd like a reply — we won't use it for anything else.

Click to browse, drag & drop, or paste a screenshot

PNG, JPG, GIF, MP4, WebM, MOV · Max 20MB each · Up to 5 files

🤖
Online · instant AI help
Broxer