T

Senior Site Reliability Engineer

Tracksuit

Auckland, New Zealand · Full Time

Be the first to apply

Experience
Any
Salary
—
Openings
1
Posted
4 days ago
Work mode
In office
Resume
Required to apply

Where you'll work

Sign in to tell us what does and doesn't work for you here — it sharpens every match we show you.

Job description

About Tracksuit

Tracksuit is a dynamic company focused on empowering marketers to demonstrate the effectiveness of their brand-building efforts. Our platform provides teams with actionable data to make informed decisions, justify budgets, and track progress without relying on cumbersome reports or exorbitant fees. Currently, we monitor over 1,000 brands across 25 countries and operate offices in Auckland, Sydney, London, and New York City. Our culture is founded on "high care, high performance," emphasizing excellence alongside mutual support.

Role Overview

We are seeking a Senior Site Reliability Engineer to join our Auckland-based team. In this role, you will oversee how reliability, observability, and security are implemented throughout our AWS infrastructure and platform layers—including agents, MCP servers, and advanced model-driven features. You will enable individual teams to efficiently operate their services by establishing golden paths, enforcing guardrails, providing tooling, and offering support during incidents.

Key Responsibilities

  • Develop and sustain cloud infrastructure and standardized paths facilitating resilience, security, and scalability by default.
  • Empower teams with self-serve observability tools encompassing monitoring, alerting, logging, tracing, and automation of provisioning and deployments.
  • Streamline incident response through tailored tooling, up-to-date runbooks, and conducting thorough post-incident reviews leading to improvements.
  • Define identity and least-privilege frameworks for non-human actors such as agents and CI systems, ensuring access with minimal risk.
  • Create clear infrastructure and operational documentation compatible with both humans and automated agents.
  • Establish standards for safe releases by embedding guardrails and pipeline checks rather than gatekeeping procedures.
  • Provide transparency and control over cloud and inference-related expenditures to allow cost management.
  • Coach engineering teams on reliable operational practices and collaborate with Engineering and Product teams to balance new features with platform stability.

Required Qualifications and Attributes

  • Proven experience leading incident command during live production issues and managing platforms through significant growth phases.
  • Advanced capabilities in infrastructure management including Infrastructure as Code (Terraform, Terragrunt, CDK), containers (ECS), cloud platforms (preferably AWS), CI/CD pipelines, and scripting languages such as Python, TypeScript, or Bash.
  • Expertise with observability tools like Datadog or equivalents, understanding distributed tracing, structured logging, and incident management principles.
  • A strong security focus, especially with networking, cloud architecture, and the security challenges of agentic systems, including secure credential handling and least privilege enforcement.
  • Experience utilizing coding agents in workflows (e.g., Claude Code), with informed perspectives on their applications and limitations.
  • Ability to navigate complex technical environments with a focus on business outcomes, effectively communicating system health and risks to diverse audiences.
  • A collaborative and kind team player attitude.
  • Bonus: Enthusiasm for marketing, design, and delivering outstanding user experiences.

Technology Stack

We utilize technologies including AWS, ECS, Terraform, Terragrunt, GitHub Actions, Datadog, Claude Code, Linear, Notion, Postgres, DynamoDB, and Snowflake.

Company Culture & Benefits

Our closely-knit, ambitious team is committed to transparency, trust, continuous learning, and improvement. Benefits include competitive salary reviewed biannually, Employee Share Option Program (ESOP), generous health and wellness perks including annual wellness bonus, premium EAP access, six weeks paid annual leave, comprehensive parental leave policies, a personal learning and development budget, plus flexible work arrangements favoring office presence but accommodating work-from-home as needed. On joining, you also receive a personalized Tracksuit kit reflecting our energetic and practical style.

Additional Information

Artificial intelligence tools may assist in aspects of the hiring process such as application review and resume analysis, but all hiring decisions are ultimately made by humans. For inquiries regarding data handling, please reach out. Tracksuit commits to transparent and fair remuneration, with regular full-cycle compensation reviews and opportunities for advancement.

Level

Senior

Tools & software

Amazon Web Services AWS required Terraform required

How they work

Communication Teamwork & Collaboration Problem Solving Work Ethic
🤖
Online · instant AI help
Broxer