Senior MLOps Engineer
Institute of Foundation Models
Abu Dhabi, United Arab Emirates · Full Time
Be the first to apply
- Experience
- 4+ yrs
- Salary
- —
- Openings
- 1
- Posted
- 1 day ago
- Work mode
- In office
- Resume
- Required to apply
Where you'll work
Sign in to tell us what does and doesn't work for you here — it sharpens every match we show you.
Job description
About the Institute of Foundation Models (IFM)
The Institute of Foundation Models is a specialized research center dedicated to the creation, analysis, deployment, and risk management of large-scale artificial intelligence systems. The institute fosters innovation in foundational AI models and their real-world applications by providing scalable infrastructure that supports research, education, and industry integration.
Role Overview
Joining the engineering team means working where machine learning intersects with systems architecture. You'll build and manage cloud infrastructure, orchestration layers, and deployment processes that power state-of-the-art intelligent applications at MBZUAI. This includes collaborating closely with leading AI researchers and engineers to deploy large language models (LLMs), voice recognition systems, and multimodal AI solutions at scale.
Responsibilities
- Architect and oversee scalable machine learning infrastructure within AWS utilizing services such as EKS (Kubernetes), EC2, RDS, S3, and IAM access controls.
- Develop and maintain Kubernetes deployment pipelines for LLM and text-to-speech (TTS) inference using tools like Helm, ArgoCD, and monitoring solutions including Prometheus and Grafana.
- Implement and enhance model serving workflows with frameworks such as vLLM, SGLang, and TensorRT to achieve high-throughput inference performance.
- Create automated CI/CD and MLOps workflows for data versioning, model validation, and deployment, leveraging GitHub Actions, Jenkins, or AWS CodePipeline.
- Integrate user interface platforms like OpenWebUI or Gradio for showcasing models and facilitating internal evaluation.
- Work collaboratively with machine learning researchers to bring models from development to production, covering technologies like ElevenLabs TTS API, Whisper ASR, and conversational LLM systems.
- Ensure observability, optimize cloud resource costs, and uphold reliability across diverse environments.
- Develop internal tooling for tasks such as dataset curation, ongoing monitoring of model performance, and retraining pipelines.
- Maintain infrastructure as code using Terraform and Helm charts to guarantee reproducibility and compliance.
- Support real-time, multimodal AI workloads spanning voice, text, and visual data across inference clusters.
Required Qualifications and Experience
- Minimum of 4 years’ professional experience in MLOps, DevOps, or cloud infrastructure engineering for machine learning applications.
- Expertise in Kubernetes cluster management and Helm chart deployment.
- Proven experience deploying ML models using vLLM, SGLang, TensorRT, or similar model serving frameworks.
- Strong command of Amazon Web Services, including EKS, EC2, S3, RDS, CloudWatch, and IAM.
- Advanced skills with Python programming, containerization with Docker, version control systems like Git, and continuous integration/continuous deployment pipelines.
- In-depth understanding of model lifecycle orchestration, data pipeline management, and monitoring tools such as Grafana, Prometheus, and Loki.
- Strong teamwork and communication abilities, capable of collaborating effectively with researchers and software developers alike.
Preferred Professional Experience
- Hands-on experience with vLLM, Kubernetes, ElevenLabs, Whisper, Gradio/OpenWebUI, or other custom hosting for TTS/ASR models.
- Familiarity with multi-GPU workload scheduling, NCCL optimization, and integration with high-performance computing clusters.
- Understanding of security best practices, cost control, and network policies within multi-tenant Kubernetes environments and Cloudflare systems.
- Previous involvement in deploying, fine-tuning, or researching large language models and foundational AI systems.
- Awareness of data governance frameworks and responsible AI deployment in research or enterprise contexts.