R

Principal Software Engineer - DevOps / Site Reliability Engineer

Riot Games

Singapore · Full Time

Be the first to apply

Experience
5+ yrs
Salary
—
Openings
1
Posted
4 days ago
Work mode
In office
Education
Bachelor's Degree in Computer Science or related field
Resume
Required to apply

Where you'll work

Sign in to tell us what does and doesn't work for you here — it sharpens every match we show you.

Job description

About Riot Games and the Role

Founded in 2006 by passionate gamers, Riot Games is renowned for creating player-centered experiences, exemplified by the massively popular League of Legends, played by over 100 million monthly users. We seek highly skilled, curious professionals who value innovation, experimentation, and breaking traditional constraints to improve how games are developed and experienced.

The AI Efficiency team develops essential platforms and tools that enable Riot employees to harness AI effectively across various creative, development, and product activities. As the team’s infrastructure grows in importance and scale, we require a seasoned engineering leader to maintain reliability, scalability, and security in production environments.

Key Responsibilities

  • Take charge of ensuring robustness, scalability, optimal performance, and the overall operational health of the AI Efficiency web platform and the tools it hosts.
  • Architect and maintain foundational infrastructure, deployment mechanisms, and operational processes that support an expanding suite of AI production services and internal tools.
  • Enhance CI/CD pipelines, release engineering, environment control, and deployment automation to facilitate safe, fast, and consistent software delivery.
  • Set and enforce production readiness criteria including monitoring, alerting, documentation, rollback, and support protocols for all new utilities.
  • Develop and manage metrics such as SLIs, SLOs, error budgets, and reliability indicators to guide engineering trade-offs and priorities balancing speed, cost, and stability.
  • Implement comprehensive observability encompassing logs, metrics, tracing, dashboards, and proactive alerting tied to application and infrastructure health.
  • Create sustainable incident response and on-call frameworks with clear escalation paths, runbooks, and communication procedures.
  • Lead or aid in diagnosing and resolving live production issues, promoting blameless postmortems and continuous improvement.
  • Automate manual operations to reduce workload, shorten detection and recovery times, and prevent recurrent failures.
  • Deploy safe rollout methods including canary releases, feature flags, health checks, rollback strategies, and controlled environment promotions.
  • Conduct capacity planning, load testing, and performance analysis to support the platform’s growing adoption.
  • Design disaster recovery, backup, failover, and resilience plans covering critical services and infrastructure components.
  • Identify and address single points of failure and systemic operational risks across all system layers and dependencies.
  • Enhance developer experience through self-service tools, reusable components, local and test environments, deployment aids, and detailed operational docs.
  • Maintain infrastructure-as-code, configuration, secrets management, and governance systems to ensure safe, auditable infrastructure changes.
  • Collaborate across engineering lifecycles to embed reliability, operability, security, and maintainability from design through deployment.
  • Troubleshoot complex issues involving web apps, APIs, distributed services, container platforms, cloud infrastructure, networking, authentication, and external dependencies.
  • Work with ML platform engineers to integrate model-serving and inference into broader standards for observability and operations.
  • Partner with infrastructure, security, IT, platform, and compliance teams to adhere to operational and security requirements.
  • Evaluate and implement AI-driven operational tools for anomaly detection, automation, remediation, and runbook execution with appropriate controls and auditability.
  • Drive technical leadership, mentoring, documentation, standardization, architecture reviews, and tooling efforts to enhance operational excellence across teams.

Required Qualifications

  • Bachelor’s in Computer Science or related field, or equivalent professional experience.
  • More than five years’ experience in roles like Site Reliability Engineering, DevOps, Infrastructure or Platform Engineering supporting production ecosystems.
  • Proficient coding and automation abilities in languages such as Python, Go, JavaScript, or TypeScript.
  • Hands-on experience designing, running, and improving cloud production systems on AWS, GCP, Azure or similar.
  • Expertise creating and managing CI/CD pipelines, release automation, and environment workflows.
  • Deep knowledge of observability concepts including metrics, logging, tracing, dashboards, and alert systems.
  • Experience in incident response, on-call duties, root cause analysis, and continuous operational improvements.
  • Proven skills in enhancing reliability, scalability, availability, and performance of distributed architectures and web platforms.
  • Familiarity with container orchestration technologies like ECS, Docker, Kubernetes or equivalents.
  • Experience with infrastructure-as-code and configuration management tools such as Terraform, Pulumi, or CloudFormation.
  • Solid Linux systems, networking, DNS, load balancing, authentication, secrets management, and cloud security fundamentals knowledge.
  • Ability to identify operational risk patterns and drive improvements spanning multiple teams or systems.
  • Strong communication skills with ability to influence cross-functional teams and maintain clarity during high-pressure scenarios.
  • Track record in providing technical leadership, mentoring peers, and establishing engineering norms.

Desired Qualifications

  • Background supporting AI/ML platforms, inference or model-serving systems, GPU workloads, or data pipelines.
  • Experience leveraging SLOs, error budgets, and reliability metrics to inform engineering choices and priorities.
  • Work designing or improving internal developer platforms, self-service infrastructure or standardized operational paths.
  • Experience creating sustainable on-call models and operational supports across multiple teams.
  • Familiar with progressive delivery techniques like canary, blue-green deployments, feature flags, and automated rollbacks.
  • Skills in performance testing, load modeling, chaos engineering, fault injection, and failure mode analyses.
  • Expertise in backup, disaster recovery, business continuity, and multi-region failover strategies.
  • Operational experience with databases, caches, message buses, object storage, API gateways, and service meshes.
  • Knowledge of security practices including access control, secrets management, vulnerability mitigation, and infrastructure hardening.
  • Ability to balance competing demands of availability, latency, velocity, efficiency, and cost at large scale.
  • Familiarity with browser automation and end-to-end testing frameworks such as Playwright for workflow validation.
  • Understanding of integrating AI-powered tools for incident diagnostics, anomaly detection, remediation, and code review automation.
  • Awareness of operational risks and controls when AI agents access source code, CI/CD, cloud, or production environments.
  • Experience establishing governance, audits, approval workflows, and quality controls for automation in operations.
  • Comfort working in environments transitioning novel tools into reliable, supported production services.

Perks and Benefits

  • Comprehensive relocation assistance.
  • Health coverage extending to you and your immediate family.
  • Flexible paid time off policies.
  • Retirement plans with company contribution matching.
  • Life insurance, parental leave, and disability benefits.
  • Access to Play Fund for gaming to enhance player understanding and community engagement.
  • Company commitment to boosting your charitable giving with matched donations.

Minimum education

Bachelor's Degree

Industry

Gaming

How they work

Communication Teamwork & Collaboration Problem Solving Leadership

Leave it if you'd like a reply — we won't use it for anything else.

Click to browse, drag & drop, or paste a screenshot

PNG, JPG, GIF, MP4, WebM, MOV · Max 20MB each · Up to 5 files

🤖
Online · instant AI help
Broxer