Senior Staff Infrastructure & Site Reliability Engineer - Datacenter AI Engineering
Riyadh, Riyadh Province, Saudi Arabia · Full Time
Be the first to apply
- Experience
- 12+ yrs
- Salary
- —
- Openings
- 1
- Posted
- 5 days ago
- Work mode
- In office
- Education
- Bachelor's or Master's degree in Engineering, Computer Science, AI/ML or related field
- Resume
- Required to apply
Where you'll work
Sign in to tell us what does and doesn't work for you here — it sharpens every match we show you.
Job description
Company Overview
Qualcomm Middle East Information Technology Company LLC is expanding its footprint in Riyadh by hiring Data Centre Engineers to bolster infrastructure supporting AI and cloud technologies. This growth aligns with Saudi Arabia's Vision 2030, focusing on digital transformation and scalable advanced connectivity solutions.
About the Role
This position centers on creating, running, and enhancing extensive AI inference systems within datacenter environments. The engineer will ensure that Qualcomm's AI infrastructure can reliably handle demanding machine learning workloads, providing scalability and high availability. The role demands solid systems and software engineering knowledge, the ability to handle complex independent tasks, and collaboration with cross-disciplinary teams including hardware, software, and ML experts.
Primary Responsibilities
- Design, deploy, and operate large-scale AI inference datacenter systems supporting critical AI functions.
- Ensure consistent reliability, scalability, and availability of Qualcomm’s AI clusters.
- Develop and maintain supportive software tools and infrastructure around AI software platforms.
- Collaborate with architecture and hardware teams to analyze and accommodate AI workload requirements.
- Build, deploy, and maintain systems that support large language models (LLM), agentic AI workflows, and AI services.
- Work alongside modeling and software teams to enhance AI100 deployment performance.
- Identify and implement optimizations for workloads running on multi-System-on-Chip (SoC) and multi-card architectures.
- Apply Site Reliability Engineering best practices including monitoring, alerting, incident response, and optimizing system performance.
- Support production machine learning systems leveraging MLOps tools and operational best practices.
- Participate in incident reviews and help maintain operational documentation to foster continuous reliability improvements.
- Create and sustain monitoring tools, dashboards, and alerting mechanisms using platforms like Prometheus, Grafana, CloudWatch, and custom telemetry.
- Develop automation to reduce operational overhead, enhance system stability, and maintain CI/CD pipelines focused on AI services and agent deployment.
- Utilize Infrastructure-as-Code tools such as Terraform and Ansible to manage infrastructure deployments.
Required Qualifications and Skills
- Extensive experience (12+ years) in SRE or related software, systems, or infrastructure engineering roles, especially within production or datacenter environments.
- Proficient with AI and deep learning workloads related to LLMs, NLP, Vision, Audio, or Recommendation systems.
- Strong understanding of ML inference concepts like batching, token streaming, and performance tuning.
- Hands-on experience with PyTorch and familiarity with other modern ML frameworks.
- Knowledge of distributed inference, checkpointing methods, and accelerator-based compute environments.
- Experience supporting AI/ML applications in live production environments, including familiarity with LLM inference pipelines.
- Strong Python programming and scripting capabilities, coupled with experience in automation and configuration management tools.
- Solid Linux operating system skills, including shell scripting, containerization, system services, and networking basics such as DNS, TLS, HTTP/gRPC.
- Experience managing cluster schedulers (e.g., Slurm) and operating highly available distributed systems.
- Expertise in monitoring and observability tools such as Prometheus, Grafana, ELK stack, or Loki, plus incident management and reliability tracking.
- Understanding of SDLC, software release procedures, and DevOps/SRE operational reliability practices.
Preferred Additional Skills
- Experience with Generative AI, Agentic AI systems, or orchestration frameworks for LLMs such as LangChain or AutoGen.
- Familiarity with ML frameworks like TensorFlow, JAX, or Ray.
- Understanding of GPU and accelerator-based systems and high-speed networking technologies like RDMA, InfiniBand, or RoCE.
- Exposure to advanced MLOps workflows or large-scale AI platform operations.
Educational Background
- Bachelor's or Master's degree in Engineering, Computer Science, AI/ML, or a related discipline is required.
Compensation and Benefits
- Competitive salary with housing and transportation allowances.
- Stock options including Restricted Stock Units (RSUs) and performance-related bonuses.
- 16 weeks of fully paid maternity leave; 6 weeks of fully paid paternity leave.
- Employee stock purchase program and child education allowance.
- Support for relocation and immigration if required.
- Comprehensive life and medical insurance.
- Reimbursement of health and recreational membership fees through the Live+Well program.
Additional Information
Qualifications for minimum experience can be met with a combination of education (Bachelor's, Master's, or PhD) and relevant software test engineering experience. Qualcomm emphasizes equal opportunity employment and provides accommodations for applicants with disabilities during the hiring process. Candidates should be prepared to comply with company policies concerning confidentiality and security. Note that unsolicited resumes from third-party recruiters or staffing agencies will not be considered.
Level
Senior
Minimum education
Master's Degree