A

Machine Learning Platform Engineer

ActAI

Singapore · Full Time

Be the first to apply

Experience
Any
Salary
Openings
1
Posted
3 days ago
Work mode
In office
Resume
Required to apply

Where you'll work

Sign in to tell us what does and doesn't work for you here — it sharpens every match we show you.

Job description

About ActAI

ActAI is focused on creating proactive AI-driven applications accessible to over 5 billion users currently using basic software like email, notes, tasks, and calendars who are not familiar with complex AI interactions. Our vision is to embed intelligence into everyday conversations, task management, and workflows with little to no prompting, delivering high reliability in long-duration workflows, persistent context retention, and effective real-world task completion, all while minimizing AI hallucinations.

About the Role

As a Machine Learning Platform Engineer at ActAI, you will be responsible for building and managing the backend infrastructure supporting our AI solutions. Your work will encompass the full lifecycle of AI models—from training, evaluation, deployment, to inference and continuous enhancement—as well as ensuring observability and operational robustness. Collaboration with AI researchers, engineers, and product teams will be key to enabling fast experimentation and reliable deployment of AI features.

Key Responsibilities

  • Develop and maintain the ML platforms and infrastructure powering ActAI's AI offerings.
  • Architect systems for training, evaluating, deploying, and inferring machine learning models effectively.
  • Enhance model serving infrastructure to support high-throughput, low-latency use cases.
  • Optimize AI system reliability, scalability, latency, and cost-efficiency.
  • Create dependable pipelines for data preparation, model training, evaluation, release, and ongoing improvements.
  • Build tools and platforms that accelerate the experimentation and delivery of AI models.
  • Implement benchmarking and evaluation frameworks to track model quality and performance.
  • Establish production-grade observability, monitoring, tracing, and alerting for AI/ML workloads.
  • Identify and resolve bottlenecks in the ML stack to consistently advance system performance.
  • Translate model and infrastructure requirements from AI engineers and product teams into production-ready systems.

Technology Stack

  • Python programming language
  • PyTorch and JAX frameworks for ML modeling
  • Large Language Model serving tools such as vLLM, SGLang, or TensorRT-LLM
  • Cloud-based infrastructure management
  • Distributed systems architecture and management
  • Machine learning/data pipeline design and workflow orchestration
  • GPU infrastructure and performance analysis tools
  • Vector databases and retrieval systems

Preferred Qualifications

  • Proven software engineering experience focused on production-grade system development.
  • Hands-on experience building machine learning infrastructure, platforms, or production ML systems.
  • Familiarity with model deployment, inference processes, evaluation techniques, and data pipelines.
  • In-depth knowledge of distributed systems and strategies for improving system reliability.
  • Proficiency in writing clean, maintainable, and scalable production-level code.
  • Ability to operate effectively within fast-paced, uncertain, and evolving environments.
  • A proactive attitude emphasizing ownership, experimentation, and ongoing improvement.

Expected Outcomes

  • Stable AI infrastructure that scales reliably with production demands.
  • Efficient workflows for model training, evaluation, deployment, and continuous enhancement.
  • Inference systems delivering low latency, high throughput, and cost-effective performance.
  • Reproducible, observable, and robust ML pipelines that are easy to maintain.
  • Rapid identification and diagnosis of model or infrastructure regressions.
  • Reusable ML infrastructure components that support multiple AI products efficiently.
  • Adaptive AI stack capable of quickly integrating new models, architectures, and inference techniques.

Tools & software

How they work

Teamwork & Collaboration Adaptability Initiative Accountability
🤖
Online · instant AI help
Broxer