AI/ML Automation Analyst
KAUST (King Abdullah University of Science and Technology)
Makkah, Makkah Province, Saudi Arabia · Full Time
Be the first to apply
- Experience
- Any
- Salary
- —
- Openings
- 1
- Posted
- 12 hours ago
- Work mode
- In office
- Education
- Bachelor's or Master's degree in Computer Science or related field
- Resume
- Required to apply
Where you'll work
Sign in to tell us what does and doesn't work for you here — it sharpens every match we show you.
Job description
About the Role
The AI/ML Automation Analyst will play a crucial role within the KSL AI Support Team at KAUST, focusing on MLOps infrastructure, container orchestration, and automation workflows tailored for supercomputing environments. This position involves creating and sustaining secure, OCI-compliant container images, robust continuous integration and delivery pipelines, and cloud-native MLOps workflows, facilitating efficient deployment and management of AI/ML tasks by researchers. The role bridges innovative Kubernetes infrastructure with the computing needs of the academic research community, emphasizing governance, enabling technologies, and community engagement.
Key Responsibilities
- Deliver prompt and effective user support through phone, walk-in assistance, email, and ticketing systems, upholding high customer service standards.
- Develop and maintain secure, HPC-compatible, OCI-compliant software container images specifically for AI/ML and data science applications.
- Architect and implement resilient MLOps workflows and pipelines optimized for supercomputing scale.
- Construct and manage CI/CD pipelines ensuring reproducible infrastructure and application deployment.
- Design, deploy, and maintain APIs for AI/ML services and inference endpoints.
- Manage Kubernetes orchestration with capabilities in CNI, CSI, and service mesh configuration and optimization.
- Oversee container registries such as Harbor and model registries including MLFlow and Kubeflow Model Registry.
- Support computational readiness and artifact governance reviews to ensure compliance with institutional standards.
- Advise users on efficient resource utilization for AI/ML workflows, ensuring security and policy adherence.
- Implement usage monitoring and reporting mechanisms.
- Conduct performance tuning and debugging for MLOps and cloud-native workflows.
- Develop benchmarks and regression tests to support system procurement and cluster stability.
- Deploy and maintain observability tools such as Prometheus, Grafana, NVIDIA DCGM, and Grafana Loki.
- Contribute to evaluation and selection processes for future infrastructure technologies.
- Create detailed training materials and documentation on MLOps platforms, Kubernetes, and containerization tools.
- Lead workshops on CI/CD, container orchestration, and best MLOps practices.
- Engage in knowledge transfer and provide individual consultations to researchers regarding automation infrastructure usage.
Qualifications and Skills
- Bachelor’s or Master’s degree in Computer Science, Data Science, Computational Science, Artificial Intelligence, or a related discipline.
- Certifications such as Certified Kubernetes Administrator (CKA), Certified Kubernetes Application Developer (CKAD), Certified Kubernetes Security Specialist (CKS), or Certified Cloud Native Platform Engineer (CNPE) are highly desirable.
- Proven experience in designing and managing complex, reliable MLOps pipelines.
- Practical expertise in API design and deployment.
- Experience developing CI/CD pipelines that guarantee reproducibility in infrastructure and workflows.
- Familiarity with academic or research computing environments is preferred.
- Strong skills in Kubernetes ecosystem components including CNI, CSI, and service mesh technologies.
- Advanced proficiency with MLOps frameworks and workflow development.
- Expertise in building secure, OCI-compliant container images.
- Programming skills in Python, with additional experience in Go and shell scripting (Bash).
- Solid Linux/Unix system administration capabilities.
- Desired knowledge includes experience with workflow orchestration tools like ArgoCD, Airflow, DASK, and Spark, ML serving platforms such as Kubeflow, KServe, and Seldon, and observability tools (Prometheus, Grafana, NVIDIA DCGM, Grafana Loki).
- Understanding of Model Context Protocol and agentic frameworks.
- Competency in deploying inference services at scale, managing container/model registries, adopting GitOps, Infrastructure as Code tools (Terraform, Ansible), and integrating HPC schedulers like SLURM with cloud environments.
Minimum education
Master's Degree
Industry
Higher Education