Senior Machine Learning Engineer
Singapore · Full Time
Be the first to apply
- Experience
- Any
- Salary
- —
- Openings
- 1
- Posted
- 3 days ago
- Work mode
- In office
- Resume
- Required to apply
Where you'll work
Sign in to tell us what does and doesn't work for you here — it sharpens every match we show you.
Job description
About the Role
We are looking for an experienced Senior Machine Learning Engineer to lead the development and maintenance of core machine learning infrastructure and subsystems. This pivotal role involves overseeing the entire ML lifecycle from conceptualizing ambiguous requirements into feasible solutions, designing scalable systems, and delivering robust ML services into production environments. The role demands hands-on expertise combined with strategic systems thinking.
Key Responsibilities
- Design, deploy, and sustain foundational ML systems that support long-term AI functionalities.
- Oversee comprehensive ML workflows including data pipeline management, model training, evaluation, inference, and continuous integration and deployment.
- Transform experimental ML research prototypes into reliable production-grade microservices.
- Diagnose, monitor, and promptly resolve complex production issues while adhering to strict latency, cost efficiency, and safety standards.
- Collaborate effectively with cross-disciplinary teams such as Research, Product, and Platform to translate ML capabilities into tangible user benefits.
- Provide architectural guidance, technical support, and mentorship to junior and mid-level machine learning engineers.
Required Qualifications
- Proven track record in deploying and managing production machine learning systems that serve active user communities.
- Expertise in software engineering principles relevant to production environments, including modularity, testing, and code maintainability.
- In-depth knowledge of modern deep learning architectures, optimization strategies, and handling edge-case scenarios.
- Self-motivated problem solver with strong communication abilities and an iterative development mindset.
Technologies Utilized
Python, PyTorch or JAX, distributed GPU training and inference pipelines.
Performance Metrics
- System Reliability: Ensure production ML services consistently meet or outperform defined performance, latency, and reliability targets.
- Operational Excellence: Rapidly identify and fix production issues minimizing service disruption.
- Business Impact: Align ML initiatives with key product metrics and company goals to drive measurable improvements.
- Engineering Standards: Elevate team quality through comprehensive code reviews and effective mentoring.
Level
Senior