Associate ML Ops Engineer
Abu Dhabi Emirate, United Arab Emirates · Full Time
Be the first to apply
- Experience
- 1–2 yrs
- Salary
- —
- Openings
- 1
- Posted
- 1 day ago
- Work mode
- In office
- Resume
- Required to apply
Sign in to tell us what does and doesn't work for you here — it sharpens every match we show you.
Job description
About AppliedAI
AppliedAI is a leading artificial intelligence technology company based in Abu Dhabi, UAE. We specialize in innovative AI solutions tailored for highly regulated sectors including healthcare, insurance, government, and financial services. Our mission is to revolutionize the future of work by automating complex, document-intensive workflows that enhance efficiency and accuracy, blending human intellect with AI capabilities.
Role Overview
We are looking for an Associate ML Ops Engineer to join our expanding team. This role bridges operations and development by ensuring platform reliability, security, and scalability. It provides an excellent chance for those in the early stages of an SRE or MLOps career to deepen their expertise alongside seasoned professionals.
Key Responsibilities
- Monitor and sustain performance and availability across production, development, and staging environments.
- Work closely with DevOps, MLOps, and Development teams to identify and resolve technical issues.
- Maintain and enhance our observability system under senior engineers' guidance.
- Improve incident response procedures and documentation.
- Assist in capacity planning and optimize system performance supporting up to 10,000 requests per minute.
- Ensure infrastructure complies with security policies and regulatory standards.
- Participate in on-call rotations during standard business hours, escalating complex issues to senior staff.
- Implement and manage service level objectives (SLOs), indicators (SLIs), and agreements (SLAs).
- Collaborate on cloud cost optimization alongside system performance and reliability improvements.
- Support continuous integration and continuous deployment pipelines and workflows.
Required Skills and Experience
- Familiarity with infrastructure best practices such as AWS and Azure Well-Architected Frameworks.
- Understanding of infrastructure design for high availability, fault tolerance, and disaster recovery.
- Knowledge of security fundamentals including least privilege principles, network segmentation, and encryption.
- Awareness of infrastructure compliance and governance frameworks.
- Interest in financial operations (FinOps) and cost management techniques.
- Exposure to Infrastructure as Code principles emphasizing modularity, reusability, and version control.
- Understanding of observability methodologies including logging, metrics, and tracing.
- Between 1 to 2 years of practical experience in Site Reliability Engineering, DevOps, or related disciplines; internships and hands-on projects are also considered.
- Experience employing AWS services such as Lambda, ECS, Fargate, ALB, ELB, API Gateway, Route53, CloudFront, AppSync, DynamoDB, RDS(PostgreSQL), Aurora, EventBridge, SNS, SQS, Security Groups, Secrets Manager, Systems Manager, IAM, ECR, CodeBuild, and CodeDeploy.
- Familiarity with monitoring and observability tools.
- Basic proficiency in infrastructure-as-code tools like CDK and Terraform.
- Insight into event-driven architectures, containerization, and microservices.
- Foundational scripting and automation capabilities to support workflows.
- Strong analytical problem-solving and structured debugging methods.
- Experience working within Agile development frameworks is advantageous.
Preferred Qualifications
- Certifications from AWS, Azure, or GCP providers.
- Experience with Next.js, Node.js, and Python programming languages.
- Familiarity with authentication platforms such as Auth0 and Single Sign-On (SSO).
- Knowledge of regulatory standards including SOC 2, HIPAA, GDPR, PCI DSS.
- Interest in machine learning and large language model (LLM) operations.
- Exposure to multi-region AWS deployment strategies.
- Inclination to work with high-traffic system environments.
- Basic database management and optimization skills.
- Understanding caching mechanisms and content delivery network (CDN) implementations.
- Knowledge about data lifecycle management and Extract, Transform, Load (ETL) processes.
- Experience with vector and graph database concepts is an asset.
What We Provide
- Hands-on opportunities with advanced technologies.
- Mentoring from experienced SRE, architecture, DevOps, and MLOps engineers.
- Collaborative teams focusing on architecture, development, DevOps, and MLOps.
- Room for career advancement in a fast-growing startup environment.
- Work alongside a distributed global team.
- Standard working hours with flexibility for emergency support when required.
- A clear progression path toward greater responsibilities in site reliability engineering.
Candidate Attributes
- Excellent verbal and written communication.
- A proactive problem-solving approach.
- Excellent teamwork and collaboration skills.
- Self-driven with initiative to learn independently.
- Comfortable adapting to a dynamic, fast-paced startup culture.
- Keen to pursue ongoing professional development.
Additional Benefits
- Work for a top-tier artificial intelligence technology company.
- Supportive, innovative workplace culture.
- Fast-paced entrepreneurial atmosphere encouraging forward-thinking.
- Professional growth and development avenues.
- Opportunity to be part of a vibrant ecosystem based at our Abu Dhabi headquarters.
- Annual leave entitlement of 21 paid days.
- Company-sponsored health insurance coverage.
- Visa sponsorship available for candidates relocating internationally.
Industry
Artificial Intelligence