- Experience
- 2+ yrs
- Salary
- —
- Openings
- 1
- Posted
- 2 days ago
- Work mode
- In office
- Education
- Bachelor's degree or equivalent experience
- Resume
- Required to apply
Where you'll work
Sign in to tell us what does and doesn't work for you here — it sharpens every match we show you.
Job description
About Crusoe
Crusoe is dedicated to advancing the abundance of energy and intelligence by creating the only vertically integrated AI infrastructure company from the ground up. Owning and operating every component from electrons to tokens, Crusoe supports the world's most ambitious AI workloads. This approach addresses the critical challenge of power limitations in AI compute by prioritizing energy efficiency and fostering faster innovation.
Role Overview
We are seeking a Staff Software Engineer to join our Cloud Availability Platform team. You will be instrumental in developing Conductor, our autonomous control plane designed to manage thousands of AI accelerators efficiently. The role involves building greenfield distributed systems that optimize power, cost, and compute in real time. You will work collaboratively with senior engineers and cross-functional teams spanning hardware, energy management, data center construction, and customer support.
Key Responsibilities
- Design, develop, and maintain distributed services and control systems that enable live AI workloads to be drained, checkpointed, replaced, and resumed without manual input.
- Collaborate closely with staff and principal engineers to review system architectures and implement safety mechanisms for autonomous operations.
- Engage daily with diverse teams across engineering and operations to align software infrastructure with physical resources.
- Lead the development of core platform capabilities including unified observability pipelines, energy-aware scheduling for compute resources, detection of performance bottlenecks, and self-qualification of hardware components to enhance overall system throughput.
Candidate Profile
- At least two years of experience building and operating backend or infrastructure systems in production, including distributed services, control planes, schedulers, or observability pipelines.
- Strong expertise in distributed systems concepts such as state reconciliation, retry mechanisms, consistency trade-offs, and live system automation.
- Proficiency in systems programming with modern languages like Go, Rust, or C++.
- Ability to navigate ambiguous problems, design clear technical solutions, and drive features from inception to deployment.
- A bias towards iterative delivery and practical production software rather than excessive documentation.
- Bachelor's degree or equivalent experience in Computer Science, Engineering, or related technical fields.
Preferred Qualifications
- Experience or strong familiarity with GPU telemetry, high-performance networking fabrics (e.g., NVLink, InfiniBand, RoCE), or hardware power and thermal characteristics.
- Background in designing scalable telemetry and observability systems spanning compute, storage, and networking.
- Knowledge of scheduler internals such as Kubernetes, Slurm, or other distributed orchestration frameworks.
- Application of statistical techniques, anomaly detection, or time-series forecasting for predictive maintenance on noisy operational datasets.
- Understanding of zero-trust security architectures, policy-based access controls, and automated multi-tenant audit streaming.
Benefits
Crusoe offers a comprehensive benefits program supporting financial security and health, including pension plans, private health and dental insurance, income protection, life assurance, and more.
Compensation and Equal Opportunity
Compensation will be structured as salary or hourly pay based on education, experience, and skill alignment with market standards. Crusoe is an equal opportunity employer committed to non-discriminatory hiring practices.
Level
Mid
Minimum education
Bachelor's Degree