R

Manager, Software Engineering - SRE Site Lead: Dublin ROC

Riot Games

Dublin, County Dublin, Ireland · Full Time

Be the first to apply

Experience
2+ yrs
Salary
Openings
1
Posted
3 days ago
Work mode
In office
Education
Bachelor's or Master's in Computer Science or related field or equivalent experience
Resume
Required to apply

Where you'll work

Sign in to tell us what does and doesn't work for you here — it sharpens every match we show you.

Job description

About the Role

The Riot Operations Center (ROC) delivers continuous 24/7 live incident response to ensure an optimal experience for Riot's players. With three globally distributed sites, ROC embraces Site Reliability Engineering (SRE) principles, prioritizing proactive incident reduction and fostering an engineering-first culture. Each site operates autonomously while aligning with the overarching ROC vision and methodologies.

As the Software Engineering Manager and Site Lead for the Dublin ROC, you will oversee high-quality incident response during your segment of the follow-the-sun model. This position involves cultivating both the skills and mindset of your engineering team to excel at technical triage, system diagnosis, and incident command. You will lead engineers to become adept systems triagers, accelerating problem localization and root cause analysis within Riot’s ecosystem.

You are a great fit if you are motivated by nurturing an evolving engineering team and mentoring engineers to advance in their careers. You understand SRE as a vital role and champion foundational engineering skills like early detection, problem identification, and structured triage. Building supportive relationships and iterative approaches to complex challenges are at the core of your management style, appreciating that difficult tasks can hold significant value.

Key Responsibilities

  • Lead and develop the Dublin ROC engineering team, focusing on growth and performance.
  • Provide mentorship to improve both technical capabilities and soft skills among engineers.
  • Oversee incident response scheduling and participate as an on-call Incident Commander.
  • Manage team capacity for both project initiatives and incident reaction, tracking and reporting site performance metrics.
  • Serve as the site’s technical authority, engaging hands-on in code and design reviews, and shaping standards for alerting and automated incident response.
  • Champion advancements toward AI-driven incident response and system triage solutions.
  • Coordinate with stakeholders during major incident investigations and critical product launches in partnership with TPM teams.
  • Handle local HR responsibilities such as recruiting and retaining engineering staff for the site.

Required Qualifications

  • Bachelor’s or Master’s degree in Computer Science or a related discipline, or equivalent professional experience.
  • Minimum of 2 years as a Senior Software Engineer or in a more senior engineering position.
  • At least 2 years of experience managing engineering teams including hiring, coaching, and career development.
  • Proven expertise in diagnosing issues in large-scale production software you did not originally develop.
  • Experience serving as Incident Commander, capable of leading high-pressure incident resolution across various roles.
  • Skill in minimizing alert fatigue through effective alert design.
  • Background in championing SRE best practices and nurturing engineers in SRE methodologies.
  • Experience designing, coding, and maintaining high-capacity, scalable, and high-performance backend systems.
  • Familiarity with containerized environments and container orchestration systems such as Marathon, Mesos, Kubernetes, GKE, or Amazon ECS.
  • Ability to collaborate across diverse teams and attain consensus on technical standards.

Preferred Qualifications

  • Over 4 years of experience in a Site Reliability Engineering role.
  • Experience operating within a global follow-the-sun team structure and effectively communicating across multiple time zones.
  • Knowledge of distributed computing architectures, especially microservices.
  • Understanding of relational database technologies like MySQL.
  • Familiarity with Continuous Integration/Continuous Deployment pipelines, including Jenkins, GitHub Actions, or similar tools.
  • Insight into software performance optimization and latency impacts relevant to online gaming.
  • Experience using AWS or comparable cloud platforms.

Our Perks

  • Flexible work arrangements supporting work/life balance with an open paid time off policy.
  • Comprehensive medical, dental, and life insurance coverage.
  • Parental leave benefits extend to employees, their spouses/domestic partners, and children.
  • Retirement support with company matching contributions.
  • Encouragement of charitable involvement with company matching of donations and volunteer time.

Additional Information

At Riot Games, our mission is to prioritize player experience in every aspect of our work. Whether crafting new player-centered features or supporting operational excellence, each team member contributes to this goal. We promote collaboration and value unique perspectives, striving to foster an environment where everyone is empowered to excel daily.

Minimum education

Bachelor's Degree

Industry

Gaming

How they work

Teamwork & Collaboration Problem Solving Leadership Relationship Building
🤖
Online · instant AI help
Broxer