- Experience
- 10+ yrs
- Salary
- —
- Openings
- 1
- Posted
- 5 days ago
- Work mode
- In office
- Resume
- Required to apply
Where you'll work
Sign in to tell us what does and doesn't work for you here — it sharpens every match we show you.
Job description
About AI71
AI71 is a leader in artificial intelligence, delivering innovative, secure, enterprise-grade applications that enable developers, businesses, and governments to address complex challenges. Focused on sector-specific needs, AI71 bridges advanced AI technologies with practical impact while prioritizing research and responsibility to create transformative results.
Role Overview
The MLOps Engineer will architect and define the machine learning infrastructure and reliability strategies across AI71's platform. This includes managing deployment, fine-tuning, and scalable serving of large language models (LLMs) and other deep learning models. The role entails decision ownership on architecture for SaaS and on-premises models, mentoring engineers, and leading multi-quarter ML infrastructure strategies.
Key Responsibilities
- Design ML infrastructure architecture: deployment strategies (using vLLM, Triton, or TGI), pipeline engineering with MLflow or Kubeflow, and cloud-native infrastructure on AWS, Azure, or GCP.
- Develop and oversee ML system reliability targeting monitoring, latency, throughput, availability, and incident management across both research and production environments.
- Provide mentorship to senior MLOps engineers and elevate operational standards across teams.
- Lead cross-team projects to enhance inference speed and reduce costs, including use of distributed training frameworks like DeepSpeed, FSDP, and Accelerate.
- Collaborate with ML researchers, product teams, and leadership to formulate long-term ML infrastructure strategies.
- Ensure ML infrastructure scalability across managed SaaS and fully air-gapped on-prem deployments.
Required Qualifications
- Over 10 years of experience in MLOps, ML infrastructure, or machine learning engineering with strong architectural ownership.
- Demonstrated expertise in designing large-scale model deployment systems, including those for LLMs.
- Extensive experience with major cloud platforms (AWS, Azure, GCP) and advanced proficiency in Python.
- Proven mentorship record that has successfully elevated engineers to independent higher-level roles.
- Strong architectural acumen for ML systems in both managed SaaS and disconnected air-gapped/on-premise environments.
- In-depth knowledge of Kubernetes at the architectural level, including GPU scheduling, multi-tenancy, operators, and managing failure modes in distributed shared clusters.
- Excellent communication, stakeholder engagement, and decision-making abilities, with a commitment to fostering diverse, inclusive engineering teams.
Preferred Additional Skills
- Experience managing production reliability at the platform level, including defining SLOs, incident command, and conducting postmortems to improve reliability across teams.
- Architectural experience with distributed training and fine-tuning at scale using frameworks like DeepSpeed, FSDP, and Megatron-LM, covering cluster design, checkpointing, and failure recovery.
- Strong GPU systems knowledge including CUDA, NCCL, interconnect topologies (NVLink, InfiniBand/RoCE), and diagnosing multi-node performance and communication issues.
- Portfolio-level model optimization strategies such as quantization (FP8, AWQ, GPTQ) and speculative decoding with measurable cost or latency benefits.
- Experience in compliance-driven architectures within regulated or security-sensitive environments, including model governance, lineage tracking, auditing, and secrets management.
- Expertise in on-premises and air-gapped ML delivery architectures at scale.
- Proven capability building and maturing MLOps practices within growing organizations, including defining standards, platform abstractions, and durable operational pathways.
- Knowledge of bare-metal GPU cluster architectures, including scheduling systems like Slurm or Kubernetes and managing hardware lifecycle in customer or owned data centers.
Desirable Additional Qualifications
- Industry thought leadership through conference speaking, technical writing, or open-source contributions to inference, serving, or ML infrastructure projects, preferably with maintainer roles.
- Experience in C/C++ or CUDA for performance-critical kernel development.
- Arabic language proficiency.
Why Join AI71
- Engage in mission-driven work on pioneering AI projects addressing critical sector challenges with a passionate team.
- Access unmatched resources and world-class models to innovate and solve impactful problems.
- Benefit from competitive compensation, comprehensive benefits, and significant prospects for career advancement as a core team member.
- Work in a flexible environment equipped with the latest technologies needed to excel.
Industry
Artificial Intelligence