D

Principal AI Ops Engineer

Discovered MENA

Abu Dhabi, United Arab Emirates · Full Time

Be the first to apply

Experience
Any
Salary
—
Openings
1
Posted
1 week ago
Work mode
In office
Resume
Required to apply

Where you'll work

Sign in to tell us what does and doesn't work for you here — it sharpens every match we show you.

Job description

About the Role

A premier organisation based in Abu Dhabi is seeking a Principal AI Ops Engineer to helm the operation of large-scale AI and large language model (LLM) systems in production environments. This role involves responsibility for inference, model serving, deployment, observability, reliability, and system performance management.

This is an individual contributor role at a senior level, requiring a fusion of deep expertise in AI infrastructure alongside solid foundations in software and reliability engineering. The focus is on setting operational standards to ensure safe, dependable, and efficient production deployment and operation of AI systems.

Key Responsibilities

  • Design and manage GPU-accelerated inference and model-serving infrastructure specifically tailored for high-demand production AI and LLM workloads.
  • Enhance model-serving processes by optimising latency, throughput, batching, quantisation, autoscaling, and infrastructure expenditure.
  • Develop automated pipelines to release models, prompts, and agent configurations with features including canary deployments, regression checks, and rollback mechanisms.
  • Implement AI-centric observability tools to monitor model layers, data retrieval, orchestration, latency, costs, and overall quality.
  • Define and uphold service level objectives (SLOs) centred on system availability, response times, and AI output quality, alongside automated alerts and incident handling protocols.
  • Oversee capacity forecasting and expenditure management for GPU and AI-related infrastructure.
  • Construct secure and scalable Kubernetes clusters and apply Infrastructure-as-Code principles for AI production deployments.
  • Create reusable deployment methodologies, tooling, and operational guidelines to empower engineering teams in consistent AI system delivery.
  • Lead resolution of complex production incidents, conduct load testing, and perform root-cause analysis on AI infrastructure and applications.
  • Provide strategic technical leadership by establishing engineering best practices for AI systems operating at scale.

Required Qualifications and Experience

  • Demonstrated track record at senior technical levels (Staff, Principal, or equivalent) operating large-scale machine learning (ML) or LLM production systems.
  • Extensive hands-on experience with GPU-powered inference and model serving frameworks such as vLLM, Text Generation Inference (TGI), TensorRT-LLM, or similar technologies.
  • Comprehensive understanding of performance trade-offs in LLM production serving including batching, quantisation, latency, throughput, and autoscaling.
  • Expertise in AI/LLM observability with ability to monitor tracing, quality metrics, drift, and regression analysis.
  • Strong foundation in reliability engineering practices encompassing SLO management, incident response, capacity planning, and post-incident reviews.
  • Proficient Python programming capabilities for creating production-grade automation and tooling relevant to infrastructure management.
  • Extensive experience with container orchestration (Kubernetes, Docker) and Infrastructure as Code tools like Terraform, preferably across multi-cloud platforms.
  • Familiarity with observability and monitoring solutions such as Langfuse, LangSmith, Arize Phoenix, Grafana, or Prometheus.
  • Ability to integrate AI evaluation and regression testing seamlessly into CI/CD workflows and production release cycles.
  • Experience working in cloud environments, notably Azure, with stringent security, data residency, and compliance conditions.
  • Knowledge of optimizing GPU and AI infrastructure costs, including FinOps practices, is advantageous.
  • Hands-on leadership style combining technical standard-setting with direct involvement in engineering and production activities.

Tools & software

Docker required Kubernetes required

How they work

Teamwork & Collaboration Problem Solving Leadership Initiative
🤖
Online · instant AI help
Broxer