Senior LLM Inference Performance & Evaluation Engineer
Singapore · Full Time
Be the first to apply
- Experience
- 5+ yrs
- Salary
- —
- Openings
- 1
- Posted
- 1 day ago
- Work mode
- In office
- Resume
- Required to apply
Where you'll work
Sign in to tell us what does and doesn't work for you here — it sharpens every match we show you.
Job description
About Bitdeer
Bitdeer is a leading global technology company specializing in Bitcoin mining and AI cloud solutions. The company offers comprehensive mining solutions, including ASIC chip design, mining rig manufacturing, and management of complex computational processes such as procurement, transport logistics, datacenter construction, and operations. With headquarters in Singapore, Bitdeer operates worldwide with a diversified 3 GW energy portfolio and deploys mining and HPC datacenters in multiple countries including the US, Bhutan, Norway, Canada, Malaysia, and Ethiopia.
Roles and Responsibilities
- Develop comprehensive benchmark pipelines measuring TTFT, inter-token latency, output throughput, request latency, concurrency curves, error rate, and GPU utilization.
- Design and implement model launch gates assessing compatibility with OpenAI/Anthropic, streaming behavior, tool integration, reasoning outputs, multimodal capabilities, and long-context scenarios.
- Maintain representative workload simulations using synthetic, replayed, and customer-like traffic patterns for various operational states such as steady and burst loads.
- Evaluate and compare different model, runtime, and provider options, producing clear recommendations for routing, fallback mechanisms, pricing, and capacity planning.
- Automate regression detection within CI/CD and staging environments to prevent silent degradations in model quality, latency, or costs due to changes.
- Collaborate with runtime engineers to identify performance bottlenecks and confirm improvements; coordinate with Site Reliability Engineers to translate benchmark results into SLOs and alert thresholds.
Required Qualifications and Skills
- Over five years of experience in ML infrastructure, performance engineering, model evaluation, QA automation, or backend testing for production environments.
- Proficient in LLM serving metrics including TTFT, TPOT/ITL, request latency, token throughput, concurrency, and GPU usage monitoring.
- Strong Python programming capabilities for automating benchmarks and evaluations; experience with Go is advantageous for integration with platform services.
- Familiarity with APIs from OpenAI and Anthropic, servers like vLLM, Dynamo, SGLang, Triton, and Kubernetes-based test setups.
- Capability to design statistically valid tests and clearly communicate tradeoffs among model quality, latency, reliability, cost, and user experience.
- Experience in developing dashboards, reports, and release gating tools for a variety of stakeholders across engineering, product, and business teams.
Work Environment and Benefits
- A culture that encourages authenticity and diverse perspectives.
- An inclusive workplace featuring open workspaces and a dynamic start-up atmosphere.
- Opportunity to grow with a rapidly expanding company and connect with experts and enthusiasts in the industrial sphere.
- Direct impact on shaping future digital asset industry developments.
- Participation in new projects and contribution to the evolution of processes and systems.
- High level of personal responsibility, autonomy, rapid development, and continuous learning opportunities.
- Attractive welfare benefits alongside professional training and mentoring programs.
Equal Opportunity Statement
Bitdeer values equal employment opportunities and complies with applicable legal standards. The company prohibits discrimination based on race, color, gender identity or expression, sexual orientation, marital or parental status, religion, political beliefs, nationality, ethnic or social origin, disability, age, indigenous status, or union affiliation.