S
Senior Solutions Architect - AI/ML Infrastructure
Remote · Full Time
Be the first to apply
- Experience
- 8+ yrs
- Salary
- GBP 185,000 – GBP 200,000 / year
- Openings
- 1
- Posted
- 1 week ago
- Work mode
- Work from home
- Resume
- Required to apply
Sign in to tell us what does and doesn't work for you here — it sharpens every match we show you.
Job description
About the Role
Join a venture-backed AI Infrastructure scale-up as a Senior Solutions Architect based in Saudi Arabia (remote). This position offers an OTE of £200,000 (comprising a £185,000 base salary plus a 10% bonus), along with equity and additional benefits. You will serve as a key technical partner for enterprise clients deploying, operating, and scaling sophisticated AI/ML workloads on a cutting-edge GPU-accelerated platform.
Key Responsibilities
- Architect comprehensive AI/ML platform solutions covering inference, training, and data processing pipelines.
- Develop reference designs for GPU cluster setup, model deployment, and multi-tenant machine learning infrastructure.
- Assess and suggest optimal inference serving frameworks such as vLLM, TGI, and Triton.
- Provide guidance on GPU fabric topologies including NVLink, InfiniBand, and RoCEv2 for distributed training environments.
- Define observability strategies integrating GPU metrics, OpenTelemetry, eBPF, and telemetry of large-scale clusters.
- Lead technical presentations, conduct workshops, and manage proof-of-concept projects.
- Serve as the primary technical advisor and escalation contact for enterprise client engagements.
- Oversee production monitoring, troubleshooting GPU usage, workload efficiency, cluster stability, and cost optimization.
- Drive root cause analyses and corrective actions for complex cross-team production challenges.
- Communicate insights to product and engineering teams to guide platform development and roadmap planning.
- Create detailed documentation for architectures and implementation procedures and mentor team members as the group expands.
Required Skills and Experience
- At least 8 years of experience in infrastructure, platform, or solutions engineering, including a minimum of 3 years specializing in AI/ML infrastructure or MLOps.
- Expertise with Kubernetes including cluster lifecycle management, workloads, operator patterns, and RBAC policies.
- Proficient with NVIDIA GPU hardware, ideally the latest generation accelerators.
- Experience with distributed training technologies such as NCCL, tensor and pipeline parallelism, and LLM inference frameworks like vLLM and TGI.
- Knowledge of GPU Operator, MIG, SR-IOV technologies, plus high-performance networking fabrics.
- Strong scripting and automation capabilities, with preference for Python, Bash, and Go.
- Hands-on familiarity with cloud platforms (AWS, Azure, GCP), including networking, identity and access management, and managed Kubernetes services.
- Competence in observability tools like Prometheus, Grafana, and OpenTelemetry.
- Excellent communication skills capable of interfacing effectively with engineering teams and executive stakeholders.
- Additional advantages include experience with Run:AI or Slurm, expertise in GPU scheduling and autoscaling, or certifications such as CKA, CKAD, or cloud Solutions Architect credentials.
Tools & software
How they work
Communication
Problem Solving
Leadership