Index Exchange

Staff Platform Site Reliability Engineer

Index Exchange

Toronto, Ontario, Canada · Full Time

Be the first to apply

Experience
8+ yrs
Salary
Openings
1
Posted
2 weeks ago
Work mode
In office
Resume
Required to apply

Where you'll work

Sign in to tell us what does and doesn't work for you here — it sharpens every match we show you.

Job description

About Index Exchange

Index Exchange is transforming digital advertising at massive scale as a leading global supply-side platform. With over two decades of pioneering experience, our proprietary technology supports top global brands and media owners, playing a vital role in maintaining an open, accessible internet. We process over 700 billion real-time auctions daily, vastly outpacing search engines in volume, by utilizing fully vertical-integrated server, network, and cloud infrastructure designed for both agility and reliability.

Role Overview

Our Cloud Platform Engineering team delivers the foundational technologies powering all facets of our operations, ranging from Kubernetes orchestration to infrastructure-as-code frameworks, container platforms, and platform APIs. Our focus is on building scalable software systems at internet scale, distinct from hands-on operational work handled by our Systems Engineering team. As a Staff Platform Engineer, you will lead architectural decisions, direct the delivery of extensive infrastructure projects, and guide the evolution of a platform that supports one of the most transaction-heavy real-time systems worldwide.

Key Responsibilities

  • Design and create a multi-tenant Kubernetes platform spanning bare-metal and public cloud environments, build scalable infrastructure-as-code frameworks, and develop essential standard libraries, SDKs, and APIs to empower all engineering teams.
  • Tackle complex distributed systems challenges including sub-millisecond real-time bidding, multi-datacenter synchronization, high-scale safe deployments, and advanced load balancing strategies.
  • Steer platform architecture by authoring RFCs, conducting design reviews, setting security and tooling standards, and making strategic decisions influencing multiple engineering domains.
  • Enhance developer velocity by constructing streamlined workflows, self-service tools, and platform APIs that enable seamless software delivery.
  • Mentor colleagues to foster technical excellence and collaborate across diverse teams including Cloud Platform Operations, SRE, Network, Security, and Software Engineering.

Candidate Qualifications

  • Minimum of 8 years experience in platform engineering, site reliability engineering, infrastructure, or DevOps roles.
  • Expertise in Linux internals including kernel tuning, networking, observability, and security.
  • In-depth knowledge of Kubernetes lifecycle management, networking, storage, RBAC, and multi-cluster deployments across bare-metal and cloud platforms (notably EKS and GKE).
  • Proficiency with infrastructure as code at scale using tools like Terraform, Ansible, and GitOps frameworks such as ArgoCD.
  • Strong programming skills in Go and/or Python for creating scalable libraries, SDKs, and platform APIs beyond scripting tasks.
  • Solid understanding of networking layers 2 through 7 including load balancing, DNS, and service discovery.
  • Proven experience influencing technical strategy across multiple teams rather than solely executing tasks within a single group.

Additional Desirable Expertise

  • Experience with distributed storage technologies such as Ceph.
  • Familiarity with big data platforms including Hadoop, Spark, HBase, and Kafka.
  • Design and implementation of observability stacks using Prometheus, Grafana, ELK, Mimir, Loki, and Tempo.
  • Management of secrets (Vault), certificate handling, and large-scale access control.
  • Expertise architecting hybrid cloud solutions combining public clouds (AWS, GCP) with on-prem infrastructure.
  • Hands-on experience operating bare-metal infrastructure across globally dispersed data centers.

Ideal Candidate Traits

  • Passionate about technology and proficient at solving complex, distributed system challenges.
  • Views operational readiness as integral to engineering, not just an afterthought.
  • Prioritizes building the correct abstractions to avoid repetitive firefighting.
  • Can coordinate alignment among cross-functional teams effectively without relying on formal titles.
  • Derives genuine satisfaction from increasing other engineers' efficiency.

Benefits and Culture

  • Comprehensive coverage for health, dental, and vision plans inclusive of dependents.
  • Flexible paid leave including health days, personal obligation days, and paid time off.
  • Competitive retirement matching contributions.
  • Equity ownership opportunities.
  • Inclusive parental leave policies covering birthing, non-birthing, and adoptive parents.
  • Annual allowances for wellness, fitness discounts, and group wellness engagement.
  • Commuter benefits and discounts where available.
  • Access to employee assistance and mental health support programs.
  • Volunteer time off yearly plus donation matching for charitable giving.
  • Regular town halls and community-driven events fostering connection.
  • Robust learning resources and programming encouraging continual professional growth.
  • A commitment to diversity, equity, and inclusion, celebrating a broad spectrum of human experiences.

Equal Opportunity and Accessibility

The company is dedicated to providing equal employment opportunities regardless of race, gender, age, disability, or any other classification. Applicants with disabilities are encouraged to request accommodations to facilitate an inclusive recruitment process.

Location

This role is based in Toronto, Ontario, Canada, at the company's corporate headquarters.

Level

Mid

How they work

Teamwork & Collaboration Problem Solving Adaptability Leadership Strategic Thinking

Leave it if you'd like a reply — we won't use it for anything else.

Click to browse, drag & drop, or paste a screenshot

PNG, JPG, GIF, MP4, WebM, MOV · Max 20MB each · Up to 5 files

🤖
Online · instant AI help
Broxer