Senior Site Reliability Engineer
Burnaby, British Columbia, Canada · Full Time
Be the first to apply
- Experience
- 5+ yrs
- Salary
- CAD 96,400 – CAD 142,660 / year
- Openings
- 1
- Posted
- 1 day ago
- Work mode
- In office
- Resume
- Required to apply
Where you'll work
Sign in to tell us what does and doesn't work for you here — it sharpens every match we show you.
Job description
About 2K
2K is a premier entertainment company known for creating iconic and influential video games such as NBA® 2K, BioShock®, Borderlands®, Mafia, Sid Meier’s Civilization®, and XCOM®, alongside popular titles like WWE® 2K, TopSpin®, and PGA TOUR® 2K. The company strives to deliver memorable and innovative gaming experiences across multiple genres via globally recognized studios including Visual Concepts, Firaxis Games, Hangar 13, Cat Daddy Games, 31st Union, Cloud Chamber, Gearbox, HB Studios, and 2K SportsLab.
At 2K, team empowerment, inclusiveness, and diversity are fundamental values that create an environment allowing employees to perform at their best.
Role Overview
The Senior Site Reliability Engineer (SRE) at 2K plays a pivotal technical leadership role, owning and overseeing production infrastructure that supports all player connections across multiple cloud providers and on-premises data centers worldwide. Responsibilities include managing 2K’s game services, account platforms, CI/CD pipelines, and developer tools with a focus on uptime, automation, and reliability during high-demand global launches and live service events.
Key Responsibilities
- Architect, build, and maintain scalable hybrid and multi-cloud infrastructure leveraging Terraform, Pulumi, and GitOps frameworks such as ArgoCD and Flux.
- Fully manage Kubernetes platforms (EKS, GKE) including cluster lifecycle, multi-tenancy, advanced networking (Istio, Cilium), and autoscaling strategies, as well as progressive delivery approaches like blue/green and canary deployments.
- Develop, operate, and optimize the full stack of monitoring and observability tools including Prometheus, Grafana, and Datadog.
- Set service level indicators (SLIs), objectives (SLOs), and error budget policies, and implement alerting that effectively reduces noise while maintaining awareness.
- Lead chaos engineering practices to proactively discover potential failures before they impact players.
- Manage incident responses and conduct post-mortem analyses focused on systemic improvements and sustainable resolutions.
- Automate repetitive tasks by creating self-service provisioning, automated remediation, and intelligent scaling solutions.
- Enhance security on platform layers through secret management tools (PasswordState, 1Password, AWS Secrets Manager) and policy-as-code frameworks such as OPA/Gatekeeper.
- Strengthen CI/CD pipelines using tools like GitHub Actions, Jenkins, and ArgoCD.
- Advocate and implement Site Reliability Engineering best practices across various 2K studios by conducting reliability reviews, authoring runbooks, and collaborating closely with development teams.
- Influence platform architecture through thorough reviews and by drafting engineering RFCs that drive technical evolution.
Required Qualifications
- At least five years of experience in Site Reliability Engineering, platform engineering, or similar infrastructure roles supporting large-scale production environments.
- In-depth expertise with Kubernetes (EKS or GKE preferred) covering aspects such as networking, multi-cluster management, and storage solutions.
- Strong proficiency in infrastructure-as-code using Terraform and/or Pulumi, with hands-on experience in Helm, Terragrunt, and GitOps toolchains (ArgoCD or GitHub Actions).
- Experience working with a mix of modern cloud platforms (AWS, GCP), virtualization (VMware), and on-premises bare metal servers.
- Knowledge of server configuration management with tools like Ansible, Puppet, and AWS Systems Manager.
- Familiarity with observability and monitoring technologies including Datadog, Prometheus, Grafana, and OpenTelemetry.
- Deep understanding of SLIs, SLOs, and error budgets along with methods to integrate these concepts operationally within engineering teams.
- Proficiency in writing production-grade software in Go, Python, or TypeScript to build automation, internal tooling, and integrations.
- Strong Linux systems expertise including internals, network protocols TCP/IP, DNS, and TLS, capable of advanced troubleshooting and debugging.
- Demonstrated leadership in incident response and post-incident evaluations focused on systemic solutions and follow-through actions.
Preferred Qualifications
- Experience supporting live-service gaming or high-traffic consumer internet applications with millions of simultaneous users.
- Advanced knowledge of service mesh technologies (Istio, Cilium) and Kubernetes networking.
- Familiarity with cloud financial operations and resource management at scale.
- Experience or interest in AI and agent-driven development.
- Relevant cloud certifications such as AWS Solutions Architect, GCP Professional Cloud Architect, CKA, or CKS.
- Background mentoring junior SREs or leading reliability-focused groups.
Compensation and Benefits
For a position located in Burnaby, Canada, the expected starting salary range is between 96400 and 142660 CAD annually. Actual base pay is influenced by the candidate's knowledge, skills, experience, and other objective factors. Eligible employees may receive additional compensation elements such as bonuses or equity awards along with a comprehensive benefits package that covers medical, financial, and other benefits. Note that temporary or intern roles are generally excluded from many compensation benefits.
Additional Information
2K is committed to providing reasonable accommodations for qualified individuals with disabilities throughout the hiring process and employment term. Requests for accommodation can be made accordingly. Candidates must be legally authorized to work in Canada without employer sponsorship; 2K does not provide visa sponsorship for this role. Communication regarding the recruitment process will only come through official company channels; beware of fraudulent contact attempts.
Level
Senior