L

Infrastructure Engineer

Letterboxd

Auckland, New Zealand · Full Time

Be the first to apply

Experience
Any
Salary
—
Openings
1
Posted
6 days ago
Work mode
In office
Resume
Required to apply

Where you'll work

Sign in to tell us what does and doesn't work for you here — it sharpens every match we show you.

Job description

About Letterboxd

Letterboxd is a platform that connects film and show enthusiasts, fueling discovery, discussions, and community through a trusted and beloved product. Millions depend on it for fast, secure, and reliable service as the user base expands, prompting the need for strong infrastructure support.

Role Overview

We are seeking an Infrastructure Engineer based in Auckland, New Zealand, reporting directly to the Head of Engineering. This role focuses on maintaining and enhancing the performance, reliability, security, and long-term sustainability of Letterboxd's technology platform. The goal is to support growth while ensuring the system remains quick, secure, recoverable, and resilient against unexpected events.

Key Responsibilities

  • Oversee the operational stability, availability, and reliability of the production platform.
  • Monitor system performance and scale capacity effectively to accommodate product and community growth.
  • Establish clear service metrics, alerts, and operational benchmarks for visibility and action on platform health.
  • Implement and manage 24/7 infrastructure support, including on-call rotations and emergency response procedures, with plans to expand the reliability engineering team.
  • Lead initiatives to strengthen infrastructure security across cloud services, networks, access controls, secrets management, and deployment environments.
  • Integrate security considerations throughout infrastructure design and implementation rather than as a final step.
  • Coordinate technical incident responses calmly and decisively.
  • Maintain and regularly test escalation, on-call, backup, and disaster recovery protocols.
  • Conduct blameless post-incident reviews to identify improvements from failures or near misses.
  • Enhance CI/CD processes, deployment tools, and release practices to ensure safe and reliable software delivery.
  • Promote the use of infrastructure as code, repeatable environments, and automation-first approaches.
  • Develop and maintain a strategic infrastructure roadmap addressing capacity, lifecycle, cost, technology updates, and phased retirement of outdated components.
  • Document key systems, decisions, and dependencies to foster shared knowledge and system durability.
  • Provide hands-on technical leadership, coaching others, and guiding sound engineering practices across teams.
  • Manage company IT resources, including hardware, security, and onboarding/offboarding processes.

Candidate Profile

Leadership in Reliability: Composed, methodical under pressure, capable of managing incidents with efficiency and minimal bureaucracy, collaborative with an emphasis on blameless improvement, and exceptionally organized with a strong sense of ownership.

Technical Proficiency: Sound judgment harmonizing speed, cost, security, reliability, and complexity, proactivity in risk identification and mitigation, preference for simple, maintainable, and automated solutions, versatile in strategic planning and practical problem-solving.

Communication Skills: Ability to simplify complex technical topics for various stakeholders, effective influence without formal authority, and produces useful documentation.

Technical Experience

  • Extensive hands-on experience managing cloud infrastructure supporting production environments for public-facing digital platforms; experience or keen interest in bare metal infrastructure is an advantage.
  • Competency in DevOps, site reliability engineering, platform engineering, or infrastructure engineering roles.
  • Proficiency working with containers, orchestration tools, infrastructure as code, and modern deployment pipelines.
  • Skill in designing monitoring, observability, alerting, backup, recovery, and incident management frameworks.
  • Strong knowledge of systems architecture, networking, security, and distributed systems.
  • Proven capability in enhancing performance, resilience, and operational maturity amidst growth.
  • Experience supporting high-traffic platforms is beneficial.
  • Passion for film, television, and related communities is a plus.

Tools Utilized

Kubernetes (preferably including bare metal), Docker, Linux/Ubuntu, Cloudflare, Grafana, Sensu, Ansible, Git, GitHub (including GitHub Actions), CI/CD tooling, and shell scripting. Equivalent experience with similar tools is valued.

Why Join Us Now?

Letterboxd is rapidly expanding — attracting more users, entering new regions, launching features, and handling increasing traffic. Our community depends on a platform that performs flawlessly. We’ve developed a highly regarded product, and now we need someone dedicated to maintaining its speed, security, and resilience while building a scalable infrastructure team and function.

How they work

Communication Teamwork & Collaboration Leadership Initiative Organisation Stress Management
🤖
Online · instant AI help
Broxer