- Experience
- 6+ yrs
- Salary
- —
- Openings
- 1
- Posted
- 1 week ago
- Work mode
- In office
- Resume
- Required to apply
Where you'll work
Sign in to tell us what does and doesn't work for you here — it sharpens every match we show you.
Job description
About Bitdeer
Bitdeer is a globally recognized technology company specializing in Bitcoin mining and artificial intelligence cloud solutions. The company designs advanced ASIC chips, manufactures mining rigs, and oversees complex operations including equipment procurement, logistics, datacenter design and construction, equipment management, along with network and facility operations. With headquarters in Singapore, Bitdeer maintains a global footprint operating a 3 GW diversified energy portfolio, deploying Bitcoin mining and HPC datacenters across multiple countries including the United States, Bhutan, Norway, Canada, Malaysia, and Ethiopia.
Key Responsibilities
- Manage and secure Kubernetes-based MaaS production environments spanning CPU platform nodes, edge ingress, and regional GPU clusters.
- Establish and maintain service level objectives, alerting systems, dashboards, runbooks, and lead incident response focused on API metrics such as availability, latency, error rate, capacity, and GPU health.
- Enhance deployment safety employing strategies like canary releases, fallback mechanisms, health-aware routing, maintenance modes, and rapid rollbacks for models, runtimes, and platforms.
- Lead capacity planning to optimize GPU usage, handle burst traffic, enforce quota and rate limits, minimize cross-region latency, and accommodate customer growth.
- Automate operational tasks with tools such as Helm, Argo CD, custom operators, scripting, and implement self-healing workflows.
- Collaborate with runtime and performance engineers to diagnose and resolve incidents ranging from public API edge issues to model worker concerns.
Candidate Profile
- Minimum six years of experience in site reliability engineering, platform engineering, or infrastructure engineering supporting production cloud services.
- Expertise in Kubernetes, including Helm, Argo CD/GitOps workflows, container networking interface (CNI) and ingress controllers, secrets management, storage, and workload scheduling.
- Hands-on experience with GPU, AI infrastructure, or high-performance computing workloads is highly advantageous.
- Proficient in monitoring and observability tools such as Prometheus, VictoriaMetrics, OpenTelemetry, and skilled in log and trace analysis to support incident resolution.
- Programming and scripting proficiency in Go, Python, and Bash; strong knowledge of Linux networking and automation in production settings.
- Demonstrated experience designing robust systems with clear service level objectives, owning operational responsibilities, and driving post-incident reviews.
Work Environment and Benefits
- A culture embracing authenticity, diversity, and varied perspectives.
- An inclusive workplace featuring open workspaces and an energetic start-up atmosphere.
- Opportunities to connect with industry pioneers and influential professionals within a rapidly growing company.
- The chance to make significant contributions shaping the future of the digital asset industry.
- Engagement in new initiatives and involvement in system and process development.
- Exposure to personal accountability, autonomy, rapid professional growth, and continuous learning.
- Access to attractive welfare benefits along with development programs including training and mentoring.
Equal Employment Opportunity Statement
Bitdeer promotes equal employment opportunities compliant with local, state, and national laws, prohibiting discrimination based on race, colour, gender identity or expression, sexual orientation, marital or parental status, religion, political views, nationality, ethnic or social background, disability, age, indigenous status, or union membership.