- Experience
- 5+ yrs
- Salary
- —
- Openings
- 1
- Posted
- 4 дня назад
- Work mode
- In office
- Resume
- Required to apply
Where you'll work
Sign in to tell us what does and doesn't work for you here — it sharpens every match we show you.
Job description
Role Overview
We are seeking an AI Engineer responsible for developing, deploying, and managing AI and large language models (LLMs) in a dual environment encompassing Google Cloud Platform (GCP) for public cloud tasks and a specialized sovereign cloud for handling classified data.
Key Responsibilities
- Construct and fine-tune machine learning and LLM models for Arabic natural language processing, document classification, computer vision/OCR, and artificial intelligence operations.
- Conduct thorough pre-deployment evaluations including accuracy baseline assessments, regression tests, safety checks, and provide justifications for GPU resource allocation.
- Improve inference performance through techniques such as quantization, task batching, and context window sizing based on usage metrics.
- Deploy AI workloads on the sovereign cloud GPU-as-a-service platform—utilizing Kubernetes, GPU partitioning on B300 nodes, resource quotas, and role-based access control mechanisms.
- Create analogous workloads on Google Cloud Platform leveraging Vertex AI and GKE with classification-based routing strategies.
- Manage the serving infrastructure including vector LLMs (vLLM), text generation inference (TGI), model version control, CI/CD pipelines, and monitoring key parameters such as latency, token usage, GPU utilization, and model drift.
- Ensure all AI models comply with data sovereignty laws and AI governance policies including ZATCA, Saudi Data & AI Authority (SDAIA) ethics standards, generative AI guidelines, and personal data protection regulations.
Candidate Requirements
- Minimum of five years of experience in machine learning or AI engineering with practical exposure to production deployment of large language models.
- Proficiency in Python programming, PyTorch framework, and Hugging Face libraries.
- Hands-on experience managing production Kubernetes clusters and GPU-accelerated inference services.
- Familiarity with GCP Vertex AI or equivalent cloud AI platforms.
Skills
Tools & software
Python
required
Kubernetes
required
PyTorch
required