R

Site Reliability Engineer

Rent The Runway

Galway, County Galway, Ireland (Hybrid) · Full Time

Be the first to apply

Experience
5+ yrs
Salary
—
Openings
1
Posted
3 hours ago
Work mode
Hybrid
Resume
Required to apply

Where you'll work

Sign in to tell us what does and doesn't work for you here — it sharpens every match we show you.

Job description

About Rent The Runway

Rent The Runway (RTR) is revolutionizing fashion by creating the world’s first Closet in the Cloud. Founded in 2009, RTR has transformed the $2.4 trillion fashion industry by offering women a sustainable, cost-effective, and joyful way to wear designer apparel and accessories through subscriptions, rentals, and ownership, leveraging hundreds of brand partners. The company operates proprietary technology and an innovative reverse logistics system, earning recognition from CNBC and Fast Company multiple times.

Galway Office

Established in April 2019, RTR’s European Technology Hub in Galway is its first international office. Situated in the historic Claddagh area, the team tackles core technology challenges influencing RTR’s growth and success across software engineering, product development, machine learning, and data science, offering diverse career opportunities.

Our Platform Engineering Team

The Platform Engineering team drives operational excellence through reliability, continuous integration, test-driven development, code reviews, and open-source contributions. They support distributed, fault-tolerant cloud systems, collaborating cross-functionally with IT, Engineering, Product, Security, Compliance, and Business units.

Job Role

We seek a Senior Site Reliability Engineer to lead initiatives in cloud infrastructure, software delivery, and observability. You will develop tooling, policies, and processes to enhance RTR’s scale and performance, lead projects, and deliver operational excellence through automation, self-service, and developer tools empowering the organization.

Key Responsibilities

  • Leverage Infrastructure-as-Code tools such as Terraform, Python scripting, Helm charts, and container orchestration platforms including Docker and Kubernetes, alongside Google Cloud Platform (GCP) services, to boost service reliability.
  • Embed software development methodologies to improve observability, alerting, tracing, automation, and self-healing mechanisms ensuring maximum platform uptime.
  • Coordinate end-to-end across platforms to support, detect, respond to, and report issues timely, escalating to appropriate teams for resolution.
  • Automate maintenance and operational tasks through continuous integration and continuous delivery (CI/CD) pipelines.

Candidate Profile

  • At least five years of hands-on experience with orchestration tools such as Kubernetes.
  • Strong passion for developing and refining CI/CD processes specifically related to cloud infrastructure.
  • Proficient in coding and scripting with Terraform, Helm, or Ansible and familiar with CI/CD tools like GitHub and Artifactory.
  • Practical experience using monitoring, alerting, and logging technologies including Splunk or GCP Monitoring.
  • Minimum three years maintaining production environments on cloud platforms like Google Cloud Platform (GCP), Amazon Web Services (AWS), or Microsoft Azure.
  • Over five years of software development experience using languages like Bash, Python, Golang, or Java.
  • Demonstrated achievements in optimizing existing systems, constructing resilient infrastructure, and automating workflows to reduce manual efforts.
  • Experienced in Agile methodologies with adherence to sprint cycles and scheduled deliveries.
  • Strong problem-solving aptitude including effective triaging and conducting root-cause analyses.
  • Excellent collaborator able to work well in team environments.
  • Proactive in advancing Site Reliability Engineering best practices across development and operations teams.
  • Willing to engage in on-call rotations to troubleshoot production issues, perform root cause analysis, and share findings with engineering and operations teams.

Benefits

  • Generous paid leave including annual leave, paid bereavement, and family sick leave to support personal and family well-being.
  • Universal paid parental leave for all parents plus flexible return-to-work programs.
  • Eligible for paid sabbatical after five years of continuous service for rest and rejuvenation.
  • Competitive pension plans supporting long-term financial security.
  • Comprehensive health, dental, and dependent care coverage from the first day of employment.
  • Frequent company events enhancing team spirit and enjoyment.
  • Hybrid work arrangement requiring 2-3 days per week at the Galway office, with up to 2 days remote work flexibility.

Equal Opportunity Statement

Rent The Runway is dedicated to fostering an inclusive workplace free from discrimination based on gender, marital status, age, disability, sexual orientation, religion, race, Traveller community membership, or any other legally protected status.

Tools & software

Docker required Kubernetes · 2 to 5 years required GitHub required Splunk Enterprise required Ansible required Terraform · 2 to 5 years required

How they work

Teamwork & Collaboration Problem Solving Initiative Work Ethic

Leave it if you'd like a reply — we won't use it for anything else.

Click to browse, drag & drop, or paste a screenshot

PNG, JPG, GIF, MP4, WebM, MOV · Max 20MB each · Up to 5 files

🤖
Online · instant AI help
Broxer