Senior LLMOps / AI Platform Engineer
Dubai, United Arab Emirates · Full Time
Be the first to apply
- Experience
- Any
- Salary
- —
- Openings
- 1
- Posted
- 2 days ago
- Work mode
- In office
- Resume
- Required to apply
Where you'll work
Sign in to tell us what does and doesn't work for you here — it sharpens every match we show you.
Job description
Company Overview
IC Markets Global is a leading Forex CFD provider catering to both active day traders, scalpers, and newcomers to the forex market. The company offers advanced trading platforms, low-latency connectivity, and superior liquidity, granting traders access to pricing historically reserved for investment banks and high net worth investors. The management team brings extensive experience across Forex, CFD, and Equity markets across Asia, Europe, and North America, enabling IC Markets Global to select top-tier technology and premium pricing sources.
Role Summary
We are searching for an experienced Senior LLMOps / AI Platform Engineer responsible for constructing, deploying, tuning, and managing scalable production-grade infrastructure for large language models (LLMs) and generative AI. This multifaceted position includes LLM inference optimization, GPU resource management, Kubernetes orchestration, cloud infrastructure maintenance, observability engineering, and overall AI platform operations.
Primary Responsibilities
- Deploy and oversee self-hosted LLM systems leveraging frameworks like vLLM, SGLang, and Ollama.
- Enhance LLM inference efficiency concerning latency, throughput, concurrency, GPU memory utilization, KV caching, and operational cost.
- Administer GPU workloads across a cluster of NVIDIA GPUs, ensuring resource maximization.
- Deploy and sustain AI service containers on Kubernetes platforms such as AWS EKS using Docker and Helm charts.
- Implement robust LLM reliability features including health monitoring, automated recovery mechanisms, and strategies for model refresh and restarting.
- Set up observability stacks employing Langfuse/LangSmith, OpenTelemetry, Prometheus, and Grafana to monitor system health and performance.
- Deploy and fine-tune Retrieval-Augmented Generation (RAG) architectures, embedding models, and vector database technologies including Qdrant and Milvus.
- Support AI agent workflows developed with LangChain and LangGraph frameworks.
- Design, build, and maintain continuous integration and continuous deployment (CI/CD) pipelines targeting AI services and related infrastructure.
- Diagnose and resolve production issues spanning LLMs, GPU performance, Kubernetes environment, networking, and AI application layers.
Required Expertise
- Proficient in Python programming and FastAPI framework for backend service development.
- Experience deploying and managing LLMs, specifically using vLLM, Hugging Face models, and self-hosted configurations.
- Strong command of container orchestration with Kubernetes, containerization via Docker, Helm package management, and cloud operations on AWS.
- In-depth knowledge of NVIDIA GPU inference processes and performance tuning.
- Familiarity with LangChain, LangGraph, and LangSmith AI toolkits.
- Expertise in building and managing RAG systems, semantic embeddings, and vector databases like Qdrant.
- Implementing LLM-focused observability and monitoring solutions.
- Solid experience with PostgreSQL and Redis databases.
- Working knowledge of GitHub Actions and general CI/CD tooling.
- Excellent troubleshooting skills for production environments involving AI models, hardware accelerators, Kubernetes, and networking.