Cygnify

Machine Learning Platform Engineer

Cygnify

Singapore · Full Time

Be the first to apply

Experience
Any
Salary
Openings
1
Posted
34 minutes ago
Work mode
In office
Resume
Required to apply

Where you'll work

Sign in to tell us what does and doesn't work for you here — it sharpens every match we show you.

Job description

About the Company

We collaborate with A1, a startup backed by BJAK, dedicated to creating advanced AI-native applications that transform communication and productivity. A1's flagship product, AI Email Triage, revolutionizes email handling by shifting user focus from composing and reading emails to supervising AI-generated content, resulting in greater efficiency and user satisfaction.

A1's technology includes Agentic AI capable of complex multi-step reasoning and external tool integration, Permission-Based Actions that require user consent before modifying calendars or sending emails, ensuring user control, and Context & Memory features that adapt to user preferences for personalized assistance.

BJAK is Southeast Asia's largest insurance platform with expanding operations in Japan, the UK, and beyond.

Role Summary

The Machine Learning Platform Engineer role involves designing and maintaining the infrastructure supporting A1's AI functionalities. You'll develop systems handling model training, testing, deployment, inference, and monitoring. Collaborating closely with AI engineers, researchers, and product teams, you'll create scalable, cost-effective, and reliable production environments. Your work will enable swift experimentation and confident deployment of AI models.

Key Responsibilities

  • Create and manage ML infrastructure and platforms for AI applications.
  • Architect systems for training, evaluating, deploying, and experimenting with models.
  • Develop and optimize inference platforms to achieve high throughput and low latency.
  • Enhance AI system reliability, scalability, responsiveness, and cost-effectiveness.
  • Construct dependable data pipelines spanning preparation to continuous model enhancement.
  • Build tools and platforms that accelerate experimentation and deployment for AI teams.
  • Implement evaluation and benchmarking frameworks to monitor model performance and detect regressions.
  • Establish production observability with monitoring, tracing, and alerting for ML workloads.
  • Identify system bottlenecks and continuously refine ML stack performance.
  • Collaborate across teams to adapt infrastructure to evolving model and product requirements.

Performance Metrics

  • Ensure AI infrastructure supports production workloads reliably at scale.
  • Enable efficient training, evaluation, deployment, and iterative improvement of models.
  • Deliver inference systems with optimal latency, throughput, reliability, and cost-efficiency.
  • Maintain ML pipelines that are reproducible, observable, maintainable, and robust.
  • Rapidly detect and diagnose model and infrastructure regressions.
  • Develop reusable ML platform components rather than duplicating efforts for each AI product.
  • Support rapid evolution of AI stack to incorporate new models, structures, and inference methods.

Candidate Profile

  • Strong software engineering background with experience in production system development.
  • Experience in building ML infrastructure, platforms, or operational machine learning systems.
  • Proficient in model deployment, inference operations, evaluation methods, and data pipeline construction.
  • Deep understanding of distributed systems and ensuring system reliability.
  • Capability to produce clean, maintainable, high-quality production code.
  • Comfortable navigating fast-paced, ambiguous environments.
  • Proactive ownership of work with openness to experimentation and continuous improvement.
  • Familiar with technologies such as Python, PyTorch or JAX, LLM and ML inference frameworks (e.g., vLLM, SGLang, TensorRT-LLM), cloud and distributed infrastructure, ML orchestration pipelines, GPU optimization tools, and vector database retrieval systems.

Tools & software

PyTorch required

How they work

Teamwork & Collaboration Adaptability Accountability

Leave it if you'd like a reply — we won't use it for anything else.

Click to browse, drag & drop, or paste a screenshot

PNG, JPG, GIF, MP4, WebM, MOV · Max 20MB each · Up to 5 files

🤖
Online · instant AI help
Broxer