Machine Learning Operations Specialist | AI Platform Engineer
Abu Dhabi Emirate, United Arab Emirates · Full Time
Be the first to apply
- Experience
- Any
- Salary
- —
- Openings
- 1
- Posted
- 2 days ago
- Work mode
- In office
- Education
- Bachelor's degree
- Resume
- Required to apply
Sign in to tell us what does and doesn't work for you here — it sharpens every match we show you.
Job description
Overview
We are seeking an adept Machine Learning Operations Specialist / AI Platform Engineer to design, operate, and enhance scalable platforms tailored for machine learning and AI workloads. This role collaborates closely with Data Scientists, ML Engineers, Software Engineers, DevOps, Cloud, and Security teams to provide dependable ML infrastructure, streamline model deployment, and oversee the entire machine learning lifecycle from development to production.
Key Responsibilities
- Design, implement, and maintain scalable infrastructure to support machine learning and AI workflows.
- Develop and manage ML platforms facilitating model development, training, deployment, monitoring, and lifecycle management.
- Create automated MLOps pipelines for model training, validation, deployment, and retraining cycles.
- Apply CI/CD and continuous integration techniques for ML models and related applications.
- Handle model version control, artifact storage, experiment tracking, and reproducibility mechanisms.
- Automate ML workflows using appropriate orchestration and pipeline technologies.
- Deploy ML models in production environments across cloud, on-premises, or hybrid infrastructures.
- Establish scalable model-serving systems capable of managing high-throughput inference loads.
- Monitor model performance, system operations, resource use, latency, reliability, and availability.
- Implement monitoring for issues like data drift, model drift, prediction quality, and overall model production performance.
- Build automated alerting and remediation pipelines for problems related to models and infrastructure.
- Collaborate with Data Scientists to transition experimental models into production-ready solutions.
- Partner with ML Engineers to enhance deployment processes, scalability, reliability, and operational productivity.
- Support GPU and high-performance computing platforms for AI workloads.
- Optimize compute power, storage solutions, network resources, and cloud assets utilized by AI platforms.
- Automate AI infrastructure provisioning and management using Infrastructure as Code practices.
- Manage containerized machine learning environments employing Docker, Kubernetes, and Helm.
- Create reusable platform components and self-service utilities supporting data science and ML teams.
- Standardize development environments for model building, testing, staging, and production deployment.
- Enforce security protocols including access controls, secrets management, encryption, and infrastructure protections.
- Support adherence to data and model governance policies throughout the ML lifecycle.
- Maintain standards for model deployment, operational procedures, runbooks, and documentation.
- Resolve issues related to model serving, platform infrastructure, pipelines, deployment, and performance with in-depth troubleshooting and root cause analyses.
- Assist cloud migration efforts, AI platform upgrades, and extensive ML projects.
- Assess emerging MLOps tools, AI infrastructure technologies, model-serving frameworks, and automation tools.
- Define and track platform reliability and operational metrics for ML and AI services.
- Manage disaster recovery, backup, high availability, and business continuity strategies for critical AI systems.
- Work alongside cybersecurity teams to ensure AI infrastructure meets security standards.
- Continuously advance automation, scalability, observability, reliability, and developer experiences across AI platforms.
Required Qualifications
- Bachelor’s degree in Computer Science, Artificial Intelligence, Machine Learning, Software Engineering, Data Engineering, or related disciplines.
- Proven experience in MLOps, Machine Learning Engineering, AI Platform Engineering, DevOps, Cloud Engineering, or similar roles.
- Deep knowledge of machine learning lifecycle management and deploying ML systems into production environments.
- Proficiency with MLOps tools like MLflow, Kubeflow, Airflow, SageMaker, Vertex AI, Azure ML, or similar.
- Experience with major cloud providers such as AWS, Microsoft Azure, or Google Cloud Platform.
- Hands-on skills with Docker, Kubernetes, Helm, and managing containerized resources.
- Strong foundation in CI/CD workflows, Git versioning, automation, and DevOps methodologies.
- Experience implementing Infrastructure as Code via Terraform, CloudFormation, Pulumi, or comparable tools.
- Programming expertise in Python, Bash, Go, or equivalent scripting languages.
- Familiarity with ML frameworks including PyTorch, TensorFlow, Scikit-learn, or others.
- Understanding of model serving mechanisms, API design, inference infrastructure, and operational ML systems.
- Knowledge of monitoring concepts for model drift, data drift, performance metrics, and system observability.
- Applied understanding of distributed systems, cloud architecture, networking, storage, and compute resources.
- Advantageous experience with GPU or accelerated computing infrastructure.
- Insight into data pipelines, feature stores, model registries, and experiment tracking platforms.
- Understanding AI security aspects such as access control, secrets management, and data protection.
- Strong diagnostic skills for troubleshooting and resolving complex ML platform issues.
- Experience with monitoring and observability tools like Prometheus, Grafana, Datadog, or equivalent.
- Comprehension of building scalable, reliable, fault-tolerant infrastructure for critical systems.
- Excellent technical writing and interpersonal communication skills.
- Proven ability to collaborate with diverse technical teams and stakeholders.
- Solid problem-solving aptitude and analytical capabilities.
- Capacity to manage multiple complex projects within fast-moving environments.
- Relevant certifications in cloud technologies, Kubernetes, AI, ML, or MLOps are highly favored.
What We Provide
- Competitive remuneration package with full employee benefits.
- Hands-on opportunities to engineer and operate cutting-edge AI and ML platforms.
- Exposure to and professional growth in MLOps, cloud services, Kubernetes, automation, and AI infrastructure.
- Access to contemporary AI platforms, cloud facilities, GPU hardware, and MLOps tooling.
- Opportunities to lead enterprise-scale AI and ML platform projects.
- Engagement in advanced machine learning and artificial intelligence initiatives.
- Support for obtaining pertinent certifications in cloud, AI, Kubernetes, and related fields.
- A collaborative engineering culture emphasizing automation, scalability, reliability, and innovation.
- Recognition for contributions enhancing ML platform reliability and operational efficiency.
- Career advancement tracks towards Senior MLOps Specialist, AI Platform Engineer, MLOps Manager, AI Infrastructure Lead, or Head of AI Platform Engineering roles.
Minimum education
Bachelor's Degree