Bitdeer (NASDAQ: BTDR)

Senior Inference Runtime Engineer

Bitdeer (NASDAQ: BTDR)

Singapore · Full Time

Be the first to apply

Experience
6+ yrs
Salary
—
Openings
1
Posted
6 days ago
Work mode
In office
Resume
Required to apply

Where you'll work

Sign in to tell us what does and doesn't work for you here — it sharpens every match we show you.

Job description

About Bitdeer

Bitdeer stands as a leading global technology enterprise specializing in Bitcoin mining and AI cloud services. The company dedicates itself to delivering comprehensive Bitcoin mining solutions, including designing cutting-edge ASIC chips and manufacturing mining rigs. Its operations span complex computational processes encompassing equipment procurement, logistics, datacenter planning and construction, equipment management, and operational oversight of networks and facilities. Bitdeer is headquartered in Singapore and runs an extensive 3 GW energy portfolio, operating Bitcoin mining and HPC datacenters in multiple countries such as the USA, Bhutan, Norway, Canada, Malaysia, and Ethiopia.

Role Overview

The Senior Inference Runtime Engineer will oversee the vital performance-driven serving layer of the MaaS platform. This position is centered on enhancing self-hosted large language models’ efficiency, cost-effectiveness, and stability by refining the runtime infrastructure behind OpenAI and Anthropic-compatible APIs.

Key Responsibilities

  • Enhance scheduling for prefill/decode processes, manage continuous batching, optimize KV cache usage, improve speculative decoding, support extended context serving, and ensure smooth streaming.
  • Configure and manage runtimes like vLLM, Dynamo, SGLang, and TensorRT-LLM to achieve optimal latency, throughput, GPU efficiency, and operational cost savings tailored to specific models.
  • Diagnose and alleviate performance bottlenecks within GPU memory systems, high-bandwidth memory pathways, NCCL/networking layers, tokenization, frontend/proxy components, and backend model workers.
  • Lead onboarding of important models by selecting suitable runtimes, defining tensor and pipeline parallelism, determining quantization parameters, setting context lengths, and establishing rollback protocols.
  • Create runtime playbooks and safe default configurations for various scenarios including reasoning, tool integration, multimodal inputs, prompt caching, and settings specific to different providers.
  • Collaborate closely with site reliability engineers and performance evaluation teams to translate benchmarking results into practical production improvements.

Qualifications and Experience

  • Minimum six years of experience in systems engineering, machine learning infrastructure, or performance-driven backend software development.
  • Practical experience working with LLM serving runtimes such as vLLM, Dynamo, SGLang, TensorRT-LLM, Text Generation Inference (TGI), or Triton.
  • Deep knowledge of GPU memory management, CUDA/NCCL fundamentals, KV cache logic, batching strategies, streaming protocols, and complexities of distributed inference architectures.
  • Proficiency in Go or Python programming languages, capable of reading runtime source code, analyzing profiling traces, and interpreting production metrics.
  • Hands-on experience managing inference services in production environments that demand stringent latency, uptime, and cost-efficiency standards.
  • Ability to convert intricate low-level performance optimizations into tangible improvements visible to end users in terms of reliability, latency, and operational margins.

Work Environment and Culture

  • An authentic culture promoting diversity of perspectives and inclusive collaboration.
  • A respectful, open-plan office atmosphere invigorated by a dynamic start-up mindset.
  • Rapid organizational growth offering opportunities to engage with pioneers and enthusiasts of the digital asset landscape.
  • Opportunity to contribute meaningfully to the evolving digital asset industry and participate in emerging projects, systems, and process evolution.
  • Emphasis on personal responsibility, autonomy, continuous learning, and professional advancement.
  • Competitive welfare benefits coupled with development programs including training and mentorship.

Equal Opportunity Commitment

Bitdeer is dedicated to providing fair employment opportunities without bias or discrimination based on race, gender identity, sexual orientation, religion, nationality, social status, disability, age, or other protected characteristics in compliance with local laws.

How they work

Teamwork & Collaboration Problem Solving Learning Agility Accountability

Leave it if you'd like a reply — we won't use it for anything else.

Click to browse, drag & drop, or paste a screenshot

PNG, JPG, GIF, MP4, WebM, MOV · Max 20MB each · Up to 5 files

🤖
Online · instant AI help
Broxer