Grafana Labs

Staff Software Engineer - Databases Site Reliability Engineering (SRE)

Grafana Labs

Remote · Full Time

Be the first to apply

Experience
8+ yrs
Salary
EUR 117,600 – EUR 141,120 / year
Openings
1
Posted
1 day ago
Work mode
Work from home
Eligibility
Candidates must be located in the UK, Sweden, Spain, Germany, or Ireland due to remote work eligibility requirements. Applications are open to individuals meeting the stated experience and skill qualifications regardless of background.
Resume
Required to apply

Sign in to tell us what does and doesn't work for you here — it sharpens every match we show you.

Job description

Company Overview

Grafana Labs is the team behind Grafana Cloud, an extensively trusted, fully managed observability platform serving over 10,000 organizations. The platform prioritizes reliability, enables quicker incident resolution, and optimizes extensive telemetry data. Rooted in open-source principles and designed for interoperability across diverse technological stacks, Grafana Cloud integrates AI into observability and vice versa, providing unified insights across all data sources. Prominent customers include Anthropic, Bloomberg, NVIDIA, Microsoft, and Salesforce. Grafana Labs operates fully remotely with staff spread across more than 40 countries and is supported by major investors such as Lightspeed Venture Partners, Sequoia Capital, and others.

Role Summary

We are seeking a Staff Software Engineer focused on Site Reliability Engineering (SRE) to contribute to supporting our top-tier Grafana Cloud customers by enhancing the reliability of our cloud database services based on Mimir, Loki, Tempo, and Pyroscope. These databases are delivered as SaaS across AWS, GCP, and Azure platforms globally.

Key Responsibilities

  • Collaborate intimately with product engineering squads using an embedded model.
  • Own and ensure production reliability for high-SLA, complex customer configurations.
  • Design and implement automation solutions that scale reliability efforts.
  • Ensure customers consistently meet defined Service Level Objectives (SLOs).
  • Define and refine per-tenant SLOs and reliability frameworks.
  • Proactively reduce SLO budget consumption to avoid recurring incidents.
  • Function as a primary escalation point and participate in on-call rotations for incident management.
  • Lead incident response when customer impact occurs and conduct thorough post-incident reviews.
  • Contribute to design documentation and conduct code reviews.
  • Influence feature design to promote scalability and operational reliability.
  • Develop automation tools that eliminate repetitive tasks and reduce toil.
  • Enhance alerting systems to improve signal quality and minimize false positives.

Daily Activities

  • Participate in regular one-on-ones with managers and teammates.
  • Review, create, and analyze SLOs, actively seeking ways to reduce SLO burn through improved monitoring, automation, and self-healing techniques.
  • Increase customers' environment observability.
  • Develop fault-tolerant architectures considering reliability at all stages of the service lifecycle.
  • Collaborate with Engineering Leadership to influence product strategy, roadmaps, and technical designs.
  • Review pull requests and assist peers with design documents.
  • Educate teammates on Site Reliability Engineering best practices and their early implementation in development.
  • Engage in Incident Response processes from investigation to resolution, including PIR documentation and communication with customers when necessary.

Candidate Profile

  • At least 8 years of professional engineering experience with over 4 years specifically in SRE, Customer Reliability Engineering (CRE), or production engineering roles, preferably with formal CRE experience.
  • Deep expertise working with Kubernetes in cloud environments such as AWS, GCP, or Azure, including infrastructure-as-code tools like Helm, Terraform, and Jsonnet.
  • Extensive experience in technical leadership, mentoring engineers, and leading teams to impactful delivery.
  • Proven ability managing multi-tenant systems in production environments.
  • Strong background in designing and implementing SLOs and understanding their operational impact.
  • Proficiency in one or more programming languages including Go, Python, or Java.
  • Solid knowledge of Linux internals, supplemented by an understanding of networking, cloud storage, and scalable system architectures.
  • Excellent analytical and problem-solving capabilities.
  • Experience contributing calmly and effectively during incident responses while producing high-quality post-incident reports.
  • Adept at evaluating performance, scaling challenges, and potential system failure modes.
  • Comfortable working autonomously within product-centric engineering teams.
  • Strong collaborator who builds productive partnerships with product engineering counterparts.
  • Intellectually curious, transparent, action-oriented, and kind in interpersonal interactions.

Compensation and Benefits

For candidates based in Ireland, the base salary range for this position is €117,600 to €141,120, varying depending on level of experience and interview assessment. Compensation packages may include equity, bonuses (where applicable), and other benefits. Please note that compensation is country-specific and will be discussed according to applicant location.

Culture and Work Environment

  • 100% remote work culture with a worldwide team emphasizing collaboration and shared goals.
  • Rapidly scaling organization providing meaningful challenges within an evolving tech landscape.
  • Open and transparent communication with frequent company-wide updates.
  • Focus on innovation supported by autonomy to promote impactful development.
  • Strong commitment to open source principles and community-driven values.
  • Empowered teams with high trust, prioritizing effective outcomes over appearances.
  • Clear career growth opportunities with defined development paths.
  • Accessible and engaged leadership committed to being visible and approachable.
  • A passionate community of colleagues dedicated to quality and mutual support.
  • In-person onboarding to facilitate smooth integration and cultural immersion.
  • Generous global annual leave policy with 30 days off per year, including 3 mandatory shutdown days for team-wide disconnection, with compliance to local laws.

Equal Opportunity Statement

Grafana Labs is committed to equal opportunity and welcomes all applicants regardless of race, gender, nationality, identity, or any legally protected characteristics. We believe diversity strengthens our organization and is foundational as we grow.

Additional Information

This role includes access to cutting-edge AI coding assistants and developer tools supported by company-funded budgets to enhance productivity. Candidates should be based in the UK, Sweden, Spain, Germany, or Ireland due to remote work eligibility requirements.

How they work

Teamwork & Collaboration Independence Learning Agility Relationship Building Integrity Interpersonal Skills

Leave it if you'd like a reply — we won't use it for anything else.

Click to browse, drag & drop, or paste a screenshot

PNG, JPG, GIF, MP4, WebM, MOV · Max 20MB each · Up to 5 files

🤖
Online · instant AI help
Broxer