Crusoe

Senior Production Engineer, Core PE

Crusoe

Dublin, County Dublin, Ireland · Full Time

Be the first to apply

Experience
5+ yrs
Salary
Openings
1
Posted
1 week ago
Work mode
In office
Resume
Required to apply

Where you'll work

Sign in to tell us what does and doesn't work for you here — it sharpens every match we show you.

Job description

About Crusoe

Crusoe is dedicated to accelerating the abundance of energy and intelligence by constructing an integrated AI infrastructure from the base hardware to operational tokens. This vertically integrated approach supports some of the most ambitious AI workloads worldwide. Joining Crusoe means becoming part of a team committed to building the future at a rapid pace.

The current era marks a major industrial revolution, with insatiable demand for AI compute capacity. As power availability becomes a bottleneck, Crusoe addresses these challenges with an energy-first perspective that improves AI infrastructure efficiency and speed.

We seek motivated team members who are eager problem solvers and opportunity seekers, who share our ambitious vision and excel in navigating less-charted paths. Our workforce includes experts in energy, manufacturing, data center construction, and cloud services, providing a fertile environment to grow your career.

Role Overview

As a Senior Production Engineer emphasizing Operational Excellence, you will be integral in maintaining the reliability, scalability, and performance of Crusoe’s GPU cloud platform tailored for next-generation AI. This position suits engineers who are passionate about addressing complex production challenges, enhancing distributed systems, and developing automation solutions that ensure seamless infrastructure operation.

You'll collaborate across teams to bolster the operational foundation and scale infrastructure supporting complex AI and HPC workloads, ensuring the highest service standards.

Key Responsibilities

  • Work closely with cross-functional teams to establish and refine service availability indicators such as SLIs and SLOs for Crusoe’s cloud platform.
  • Engage in incident management, troubleshoot service interruptions, and participate in post-incident analysis to identify root causes.
  • Design, operate, and enhance observability frameworks using tools like Prometheus, Grafana, Alertmanager, and OpenTelemetry.
  • Detect risks affecting reliability, identify performance limitations, and recognize early signals of production problems within distributed systems.
  • Create automation and tooling aimed at minimizing manual operational work, accelerating recovery procedures, and enabling infrastructure self-healing.
  • Partner with teams in compute, networking, storage, and platform areas to improve service resilience and disaster recovery strategies.
  • Contribute to refining operational processes, fostering knowledge exchange, and promoting best practices in reliability across engineering functions.
  • Engage in ongoing professional development through mentoring and hands-on management of large-scale AI infrastructure.

Required Qualifications

  • A minimum of five years’ experience in Production Engineering, Site Reliability Engineering, or managing large-scale infrastructure operations.
  • Proven experience supporting GPU-intensive workloads, HPC settings, or latency and throughput-sensitive distributed systems.
  • Strong command of Linux/Unix operating systems, including complex kernel and user space troubleshooting.
  • Experience building or administering infrastructures focused on compute, storage, or network functionalities.
  • Knowledge of contemporary cloud infrastructure concepts such as Kubernetes, virtualization, distributed systems, and platforms like AWS or GCP.
  • Understanding of incident management methodologies and reliability frameworks like SRE or ITIL.
  • Familiarity or eagerness to deepen expertise in monitoring solutions, particularly Prometheus and Grafana.
  • Experience with infrastructure-as-code and configuration management tools such as Terraform or Ansible.
  • Proficient with scripting or programming languages including Go, Python, C, or C++.
  • Excellent communication skills and the ability to collaborate effectively within diverse engineering teams.
  • Capacity to maintain composure and efficiency when resolving complex issues in critical production environments.
  • A growth-oriented attitude with keen interest in enhancing reliability engineering, automation, and operational excellence.

Preferred Skills and Experiences

  • Hands-on experience with Kubernetes or container orchestration at scale.
  • Knowledge of change management, operational readiness assessments, or structured root cause analyses.
  • Background in designing automated healing systems or event-driven operational tools.
  • Interest and experience in scaling AI or HPC infrastructure to tackle reliability challenges in GPU-focused environments.
  • A passion for mentoring and continuous professional development within Production Engineering.

Benefits and Compensation

Crusoe provides a competitive benefits package supporting financial security, health, and well-being, including pension schemes, private health and dental insurance, income protection, and life assurance.

Compensation will be determined based on education, experience, and skillset, and may be offered as salary or hourly pay, aligned with internal and market benchmarks.

Equal Opportunity Employment

Crusoe is committed to fostering an inclusive workplace and makes hiring decisions without discrimination based on race, religion, disability, pregnancy, citizenship, marital status, gender, sexual orientation, age, veteran status, national origin, or any other legally protected status.

Tools & software

Kubernetes required Prometheus required Grafana Labs Grafana Cloud required Terraform required

How they work

Communication Teamwork & Collaboration Problem Solving Learning Agility Stress Management
🤖
Online · instant AI help
Broxer