- Experience
- 4+ yrs
- Salary
- SGD 196,000 – SGD 262,000 / year
- Openings
- 1
- Posted
- 1 day ago
- Work mode
- In office
- Resume
- Required to apply
Where you'll work
Sign in to tell us what does and doesn't work for you here — it sharpens every match we show you.
Job description
About CoreWeave
CoreWeave is a pioneering cloud platform explicitly designed for AI applications, delivering cutting-edge infrastructure and tools for innovators across AI labs, startups, and global corporations. Established in 2017 and publicly traded since March 2025, CoreWeave excels at combining high-performance infrastructure with in-depth technical proficiency to expedite AI breakthroughs.
Role Overview
As the Production Engineer – Team Lead, you will provide senior-level expertise in ensuring the stability, availability, and resilience of CoreWeave’s cloud platform, specifically supporting expansive multi-tenant GPU workloads. This role involves leading the incident command during significant service outages, coordinating cross-departmental teams, overseeing root cause analyses, and enhancing post-incident review procedures.
Responsibilities
- Provide strategic technical leadership across the cloud platform operations.
- Act as primary Incident Commander during major disruptions, managing coordination and communication.
- Define and monitor Service Level Objectives (SLOs) to maintain operational standards.
- Drive automation efforts to minimize mean time to detection (MTTD) and recovery (MTTR).
- Lead comprehensive investigations into complex distributed system incidents.
- Develop scalable incident runbooks and foster a culture of continuous improvement and reliability.
- Mentor junior Production Engineers, enhancing team proficiency and operational practices.
Candidate Profile
- Minimum 4 years of experience in production engineering, cloud operations, Site Reliability Engineering (SRE), or incident management.
- Expertise in cloud architectures, containerization, and Kubernetes infrastructure.
- Proficient knowledge of incident management frameworks, ITIL, and SRE methodologies.
- Experience with observability tools such as Prometheus and Grafana and telemetry best practices.
- Skilled in scripting and automation using Python, Bash, and Terraform.
- Strong leadership skills in high-pressure incident command and stakeholder communication.
- Track record of guiding and mentoring technical teams to develop operational excellence.
Preferred Qualifications
- Hands-on experience as Incident Commander managing high-severity service restorations.
- Advanced understanding of Kubernetes internals, distributed systems, and container runtimes.
- Practical experience creating automated, self-healing infrastructure and incident response automations.
Culture and Values
- A proactive leader who takes charge during critical incidents and drives accountable, blameless recovery.
- Analytical mindset focused on diagnosing complex distributed failures and eliminating repetitive manual tasks.
- Expertise in establishing reliability standards, SLO frameworks, and mentoring SRE teams to enhance cloud platform resilience.
- Embraces CoreWeave’s values: curiosity, ownership, employee empowerment, client focus, and teamwork.
Why Join CoreWeave?
CoreWeave offers a dynamic and fast-paced environment amid rapid growth. The company promotes independent thinking, collaboration, and innovative problem solving. You will engage with industry-leading talent and be instrumental in shaping a leading AI cloud infrastructure platform.
Compensation & Benefits
The annual base salary ranges from SGD 196,000 to SGD 262,000, with starting pay based on role-relevant skills, experience, and market factors. The comprehensive total rewards package includes a discretionary bonus, equity awards, and a broad benefits plan contingent on eligibility.