- Experience
- 5+ yrs
- Salary
- —
- Openings
- 1
- Posted
- 12 hours ago
- Work mode
- In office
- Resume
- Required to apply
Where you'll work
Sign in to tell us what does and doesn't work for you here — it sharpens every match we show you.
Job description
About the Role
We are seeking a Senior LLMOps Engineer to join our Singapore team at PatSnap. This is a technical leadership position focused on developing and maintaining robust infrastructure and platforms that support model training, inference services, and training data management. The role requires strong software engineering expertise and infrastructure knowledge, playing a key part in driving our AI transformation initiatives.
Responsibilities
- Lead development of platforms and workflows for model training, including resource scheduling, job supervision, and troubleshooting training and fine-tuning tasks.
- Create and maintain model inference systems encompassing deployment of open-source models and integration with external Large Language Model APIs, managing access controls, quotas, and requests centrally.
- Build and handle training data management systems providing dataset storage, version control, quality assurance, access regulation, and traceability features.
- Develop automation and supporting platforms for managing the lifecycle, deployment, and continuous monitoring of models, training jobs, and inference services.
- Optimize GPU and computing resource usage to enhance reliability, performance, and cost-efficiency of training and inference workloads.
- Engage in a team-based on-call rotation to deliver 24/7 operational support, promptly resolve incidents, and encourage ongoing improvements.
- Collaborate closely with AI/ML teams, engineering groups, and security personnel to deliver platform assistance and expert guidance, facilitating AI adoption company-wide.
Requirements
- Minimum five years’ experience in software development, platform engineering, Site Reliability Engineering (SRE), or MLOps with direct involvement in model training or inference infrastructure.
- Strong proficiency in Python, capable of independently creating and maintaining platforms and automation tools.
- Hands-on experience with model training workflows, fine-tuning, inference implementation, deployment of open-source models, and integrating external LLM APIs.
- Working knowledge of Linux, Kubernetes, GPU resource management, continuous integration/delivery (CI/CD), and observability tools.
- Understanding of storage solutions, version control practices, and access management for training datasets and model artifacts.
- Skilled in leveraging AI technologies to boost engineering productivity, with excellent technical judgment, problem-solving aptitude, and collaboration skills.
- Fluent in English; proficiency in Chinese is a plus.
Preferred Qualifications
- Active contributions to open-source AI infrastructure, model training, or inference tooling projects.
- Experience with distributed training strategies, optimizing inference performance, or managing GPU clusters.
Skills
Tools & software
Kubernetes
required
How they work
Teamwork & Collaboration
Problem Solving
Leadership
Initiative
Work Ethic