KAUST (King Abdullah University of Science and Technology)

AI/ML Automation Analyst

KAUST (King Abdullah University of Science and Technology)

Makkah, Makkah Province, Saudi Arabia · Full Time

Be the first to apply

Experience
Any
Salary
Openings
1
Posted
34 minutes ago
Work mode
In office
Education
Bachelor's or Master's degree
Resume
Required to apply

Where you'll work

Sign in to tell us what does and doesn't work for you here — it sharpens every match we show you.

Job description

About the Role

The AI/ML Automation Analyst will play a crucial role within the KSL AI Support Team, specializing in the development and maintenance of MLOps infrastructure and automating workflows tailored for supercomputing environments. This position emphasizes the creation of secure, OCI-compliant container images, along with designing and sustaining CI/CD pipelines and cloud-native MLOps workflows that empower researchers to deploy AI/ML workloads efficiently. The incumbent will also function as a liaison between the advanced Kubernetes-based platforms and the diverse requirements of academic researchers, contributing actively to governance, technology enablement, and cultivating the research community.

Key Responsibilities

  • Deliver responsive user support through various channels, including phone, in-person, email, and ticketing systems, ensuring questions and issues are handled efficiently.
  • Develop and maintain high-quality, secure, HPC-compatible AI/ML and data science container images adhering to OCI standards.
  • Create and manage robust MLOps workflows and scalable pipelines suited for supercomputing usage.
  • Build and sustain reproducible infrastructure deployments via CI/CD pipelines.
  • Design and implement APIs for AI/ML service endpoints and inference availability.
  • Manage Kubernetes orchestration components such as CNI, CSI, and service mesh setups, and optimize their performance.
  • Operate and maintain container and model registries including Harbor, MLFlow, and Kubeflow Model Registry.
  • Support governance through AI research computational readiness and model artifact compliance reviews.
  • Advise users on efficient resource utilization for AI/ML and MLOps workflows while ensuring adherence to security policies across the container images and workflows.
  • Implement monitoring and reporting systems for usage oversight.
  • Conduct performance tuning and debugging for MLOps and cloud-native workflows.
  • Develop benchmark tests for AI/ML workloads and maintain regression testing capabilities for current cluster environments.
  • Implement observability tools like Prometheus, Grafana, NVIDIA DCGM, and Grafana Loki to monitor resources and system health.
  • Engage in technology evaluation efforts to guide infrastructure investment decisions.
  • Prepare comprehensive training materials and deliver workshops related to MLOps platforms, Kubernetes, containerization, CI/CD pipelines, and best practices.
  • Maintain thorough documentation supporting automation tools, workflows, and conduct knowledge transfer to the KAUST research community.
  • Provide personalized consultation for researchers to optimize automation infrastructure usage.

Qualifications & Experience

  • Possess a bachelor's or master's degree in Computer Science, Data Science, Computational Science, Artificial Intelligence, or similar disciplines.
  • Certifications such as Certified Kubernetes Administrator (CKA), Certified Kubernetes Application Developer (CKAD), Certified Kubernetes Security Specialist (CKS), or Certified Cloud Native Platform Engineer (CNPE) are highly regarded.
  • Proven track record building and managing complex MLOps pipelines.
  • Hands-on API design and deployment experience.
  • Experience with creating portable and secure CI/CD pipelines for deployment reproducibility.
  • Background supporting researchers or experience within academic and research computing environments preferred.

Technical Expertise

  • Advanced knowledge of Kubernetes, including Container Network Interface (CNI), Container Storage Interface (CSI), and service mesh technologies.
  • Competency in developing, deploying, and maintaining MLOps pipelines and workflows.
  • Proficiency in CI/CD pipeline construction focused on infrastructure and application deployments.
  • Ability to create secure, OCI-compliant container images for AI/ML and data science applications.
  • API development and deployment skills.
  • Strong programming abilities in Python, with additional scripting in Go and Bash.
  • Expertise in Linux/Unix systems administration.

Desirable Skills

  • Experience using workflow orchestration tools such as ArgoCD, Airflow, DASK, and Spark.
  • Familiarity with ML serving and pipelines frameworks like Kubeflow, KServe, and Seldon.
  • Competence in managing observability stacks including Prometheus, Grafana, NVIDIA DCGM, and Grafana Loki.
  • Knowledge of Model Context Protocol (MCP) and agentic frameworks.
  • Experience scaling deployment of inference services.
  • Administration and operation of container registries (Harbor) and model registries (MLFlow, Kubeflow Model Registry, Artifact Hub).
  • Application of GitOps methodologies and Infrastructure as Code tools such as Terraform and Ansible.
  • Exposure to HPC schedulers like SLURM and strategies for HPC integration with cloud resources.

Minimum education

Bachelor's Degree

Tools & software

Kubernetes required

Leave it if you'd like a reply — we won't use it for anything else.

Click to browse, drag & drop, or paste a screenshot

PNG, JPG, GIF, MP4, WebM, MOV · Max 20MB each · Up to 5 files

🤖
Online · instant AI help
Broxer