Staff AI Scheduling & Orchestration Engineer
Singapore · Full Time
Be the first to apply
- Experience
- 6+ yrs
- Salary
- —
- Openings
- 1
- Posted
- 1 day ago
- Work mode
- In office
- Education
- Bachelor's or Master's degree in Computer Science, Electrical Engineering or related fields
- Resume
- Required to apply
Where you'll work
Sign in to tell us what does and doesn't work for you here — it sharpens every match we show you.
Job description
About Bitdeer
Bitdeer is a global leader in Bitcoin mining technology and AI cloud solutions. The company provides comprehensive services across the Bitcoin mining value chain, including ASIC chip design, mining rig manufacturing, equipment procurement, logistics, datacenter design, and operations. Headquartered in Singapore, Bitdeer operates worldwide with a diversified 3 GW energy portfolio and maintains datacenters in the United States, Bhutan, Norway, Canada, Malaysia, and Ethiopia.
Key Responsibilities
- Develop and implement sophisticated batch scheduling systems using frameworks such as Volcano and YuniKorn to enable multi-node gang scheduling.
- Manage cluster-wide admission controls and advanced job queueing with Kueue to handle large-scale AI workload traffic.
- Utilize Kubernetes Dynamic Resource Allocation and custom scheduler plugins for managing complex accelerator resource requests.
- Design topology-aware pod placement strategies that prioritize low-latency communication through NVLink and InfiniBand interconnects.
- Implement GPU sharing methods like MIG and time-slicing, alongside multi-tenancy isolation policies to improve overall cluster utilization.
- Collaborate closely with GPU Systems and Storage teams to integrate scheduling features with hardware and storage I/O operations.
- Enhance the scheduling system's reliability and scalability by addressing resource contention and deadlocks in HPC settings.
- Mentor junior engineers and lead design reviews to uphold architectural standards within the orchestration layer.
Qualifications and Experience
- Bachelor’s or Master’s degree in Computer Science, Electrical Engineering, or a related discipline.
- At least six years of experience in distributed systems engineering, specializing in Kubernetes scheduling frameworks and orchestrators.
- Deep understanding of AI workload execution and distributed training frameworks such as PyTorch Distributed, Ray, and MPI.
- Proven expertise in operating, troubleshooting, and scaling scheduling infrastructures in HPC or large-scale cloud environments.
- Strong familiarity with GPU architectures and the scheduling complexities involved in distributed AI training and inference.
- Experience with infrastructure automation and infrastructure-as-code tools, including Terraform and Go-based Operators.
- Excellent communication and leadership capabilities, with experience influencing cross-functional teams and aligning on architecture strategies.
- Capability to thrive in fast-paced engineering environments and transform complex requirements into scalable solutions.
Work Environment and Benefits
- A culture embracing authenticity and diverse backgrounds.
- An inclusive work atmosphere with open office spaces and energetic startup dynamics.
- Opportunities to engage with industry pioneers and innovators.
- Ability to directly influence the future of the digital asset industry through impactful contributions.
- Participation in new initiatives and development of systems and processes.
- High personal accountability and autonomy fostering rapid growth and learning.
- Attractive welfare benefits along with training and mentoring opportunities.
Diversity and Inclusion Commitment
Bitdeer pledges equal employment opportunities in compliance with all relevant local laws. The company prohibits discrimination based on race, color, gender identity or expression, sexual orientation, marital or parental status, religion, political views, nationality, ethnic or social origin, social status, disability, age, indigenous status, or union membership.
Level
Mid
Minimum education
Bachelor's Degree