CoreWeave

Senior Software Engineer, Cluster Orchestration (SUNK)

CoreWeave

Singapore · Full Time

Be the first to apply

Experience
8+ yrs
Salary
—
Openings
1
Posted
1 week ago
Work mode
In office
Education
Bachelor's degree in Computer Science, Engineering, or related field
Resume
Required to apply

Where you'll work

Sign in to tell us what does and doesn't work for you here — it sharpens every match we show you.

Job description

About CoreWeave

CoreWeave is the Essential Cloud for AI, designed by pioneers to empower innovators through superior infrastructure, comprehensive tools, and expert teams. Supporting leading AI labs and global enterprises, CoreWeave accelerates AI breakthroughs with cutting-edge performance. Established in 2017 and publicly listed on Nasdaq (CRWV) since 2025, it remains at the forefront of AI cloud solutions.

Role Overview

The Cluster Orchestration team at CoreWeave develops and manages Kubernetes-native platforms that enable efficient AI training and inference at scale. This includes advancing orchestration innovations such as SUNK (Slurm on Kubernetes) to ensure smooth, reliable, and optimal operation of vast GPU clusters.

In your role as Senior Software Engineer specializing in Cluster Orchestration with SUNK, you will lead technical strategy, define platform architecture, and guide critical infrastructure development. This position involves ownership of key orchestration components, managing multi-tenant scheduling, enforcing quotas, scaling systems hyperscale, and setting organizational standards for reliability and observability. You will also mentor senior engineering staff and steer vital cross-team initiatives.

Key Qualifications

  • A bachelor’s degree in Computer Science, Engineering, or an equivalent field or experience.
  • Minimum eight years of experience designing, deploying, and scaling distributed systems in production.
  • Expertise in Go programming and distributed systems architecture.
  • In-depth knowledge of Kubernetes internals, Slurm schedulers, or cloud-native technologies.
  • Proven leadership in establishing technical directions and influencing architecture across multiple teams.
  • Strong mentorship capabilities to elevate engineering practices and review intricate technical designs.

Preferred Experience

  • Familiar with orchestration and workflow tools such as Ray, Kubeflow, Kueue, Istio, Knative, or Argo Workflows.
  • Experience with distributed GPU cloud workloads or machine learning pipelines.
  • Understanding of advanced scheduling methodologies, including quota enforcement and resource pre-emption.
  • Knowledge of enterprise reliability practices like SLO definition, telemetry, and incident reviews.
  • Hands-on involvement with large-scale AI infrastructure workloads, including machine learning training or inference and high-performance computing.

Candidate Attributes

  • Passionate about defining and evolving global hyperscale cloud platform architectures.
  • Curious to explore pioneering orchestration innovations beyond SUNK for next-gen AI workloads.
  • Skilled leader in mentorship and resolving infrastructure performance, cost, and reliability challenges.

Work Culture

CoreWeave thrives in a fast-paced hyper-growth environment with a culture of curiosity, ownership, empowerment, client focus, and collaboration. Independent thinking and entrepreneurial spirit are encouraged, with ample growth opportunities and learning from industry leaders.

Compensation

Compensation includes base salary, equity, flexible vacation, and comprehensive benefits. Salary offers consider experience, skills, and market location to ensure fairness and alignment.

Minimum education

Bachelor's Degree

How they work

Teamwork & Collaboration Leadership Strategic Thinking Stress Management
🤖
Online · instant AI help
Broxer