R

Senior Observability Engineer

Rent The Runway

Galway, County Galway, Ireland (Hybrid) · Full Time

Be the first to apply

Experience
5+ yrs
Salary
—
Openings
1
Posted
3 days ago
Work mode
Hybrid
Resume
Required to apply

Where you'll work

Sign in to tell us what does and doesn't work for you here — it sharpens every match we show you.

Job description

About Rent the Runway

Founded in 2009, Rent the Runway (RTR) is revolutionizing fashion by offering the world’s first Closet in the Cloud. It disrupts the $2.4 trillion fashion sector by providing women with an innovative, sustainable, and cost-effective way to dress well every day. RTR offers access to designer apparel and accessories from hundreds of brands via customizable subscriptions, rentals, or direct ownership. The company has developed proprietary technology and unique reverse logistics operations and has been recognized multiple times by CNBC and Fast Company for innovation.

RTR’s European Technology Hub, established in 2019 in Galway, Ireland, is crucial for the company’s growth. This hub supports software engineering, product development, machine learning engineering, and data science initiatives. The Galway team works with continuous integration, test-driven development, peer reviews, pair programming, and contributes to open-source projects.

Role Overview

As a Senior Observability Engineer, you will lead the design and expansion of telemetry systems that ensure our platforms remain reliable, efficient, and resilient. You will guide observability strategies company-wide to improve incident handling, system analytics, and delivery results.

Your role involves defining metrics, alerts, and responses, enabling engineering teams to confidently build and maintain services. As a senior leader, you will own major projects, drive the adoption of observability standards, and collaborate widely to create scalable, efficient solutions.

Key Responsibilities

  • Architect, implement, and enhance observability tools using platforms such as Splunk Observability Cloud and Google Cloud Observability.
  • Develop scalable, automated telemetry pipelines utilizing Infrastructure-as-Code methods with Terraform to promote reliability and self-service adoption.
  • Establish and refine best practices for metrics, traces, logs, and event handling to generate meaningful alerts and standardized instrumentation.
  • Partner with application, platform, security, and compliance teams to integrate observability across development life cycles, including instrumentation, Service Level Objectives (SLOs), and incident post-mortems.
  • Lead the creation and application of internal guidelines for service-level indicators (SLIs), error budgets, and overall system health monitoring.
  • Promote AI-assisted development and debugging techniques within observability tools to speed up incident resolution and root cause identification.
  • Address cross-team observability challenges through simplification, standardization, and durable design solutions.
  • Participate in site reliability engineering (SRE) and platform on-call rotations to remain engaged with production issues and continuously improve alerting and tools.
  • Provide technical mentorship to engineers, helping them apply telemetry best practices for enhanced system insight.
  • Cultivate a reliability-focused culture by creating reusable frameworks, shared knowledge bases, and training programs that improve observability efficiency without increasing workload.

Candidate Profile

  • At least 5 years’ experience in site reliability engineering, DevOps, or platform engineering with a specialization in observability and telemetry design.
  • Recognized expert with deep knowledge of observability tools and practices, especially in cloud-native and distributed systems.
  • Successful experience delivering impactful projects that enhance system visibility, reduce operational risk, or improve incident response at scale.
  • Strong grasp of metrics, logging, tracing, and event management applied to both infrastructure and product services.
  • Hands-on experience managing telemetry pipelines with Terraform, continuous integration/delivery tools, and instrumentation in Kubernetes environments.
  • Solid understanding of system design concepts including service meshes, asynchronous messaging, and distributed system behavior.
  • Excellent communication skills capable of translating complex technical concepts into clear insights for technical and non-technical stakeholders.
  • Demonstrates ownership and curiosity by proactively identifying failure modes, simplifying complex systems, and leading structured root cause analyses.
  • Effective collaborator who builds consensus and leads by influence across multiple domains.

Benefits

  • Generous paid time off including annual leave, bereavement, and family sick leave, supporting employee and family wellness.
  • Universal paid parental leave for both parents and a flexible return-to-work program.
  • Paid sabbatical after five years of service to recharge and engage in personal pursuits.
  • Competitive stakeholder pension plans providing future financial security.
  • Comprehensive health, dental, and dependent care coverage starting from day one of employment.
  • Regular company-wide events and outings to foster team spirit and enjoyment.
  • Hybrid work model requiring 2-3 days per week in the Galway, Ireland office with up to 2 days remote work.

Additional Information

Rent the Runway is an equal opportunity employer committed to diversity and inclusion. Discrimination on any legally protected basis including gender, marital status, age, disability, sexual orientation, race, religion, or membership in the Traveller community is prohibited.

Tools & software

Kubernetes · 5 to 8 years required Terraform · 5 to 8 years required

How they work

Communication Teamwork & Collaboration Problem Solving Learning Agility Accountability

Leave it if you'd like a reply — we won't use it for anything else.

Click to browse, drag & drop, or paste a screenshot

PNG, JPG, GIF, MP4, WebM, MOV · Max 20MB each · Up to 5 files

🤖
Online · instant AI help
Broxer