S

Senior Solutions Architect - AI/ML Infrastructure

Scout Global

Remote · Full Time

Be the first to apply

Experience
8+ yrs
Salary
GBP 185,000 – GBP 200,000 / year
Openings
1
Posted
1 week ago
Work mode
Work from home
Resume
Required to apply

Sign in to tell us what does and doesn't work for you here — it sharpens every match we show you.

Job description

About the Role

Join a venture-backed AI Infrastructure scale-up as a Senior Solutions Architect based in Saudi Arabia (remote). This position offers an OTE of £200,000 (comprising a £185,000 base salary plus a 10% bonus), along with equity and additional benefits. You will serve as a key technical partner for enterprise clients deploying, operating, and scaling sophisticated AI/ML workloads on a cutting-edge GPU-accelerated platform.

Key Responsibilities

  • Architect comprehensive AI/ML platform solutions covering inference, training, and data processing pipelines.
  • Develop reference designs for GPU cluster setup, model deployment, and multi-tenant machine learning infrastructure.
  • Assess and suggest optimal inference serving frameworks such as vLLM, TGI, and Triton.
  • Provide guidance on GPU fabric topologies including NVLink, InfiniBand, and RoCEv2 for distributed training environments.
  • Define observability strategies integrating GPU metrics, OpenTelemetry, eBPF, and telemetry of large-scale clusters.
  • Lead technical presentations, conduct workshops, and manage proof-of-concept projects.
  • Serve as the primary technical advisor and escalation contact for enterprise client engagements.
  • Oversee production monitoring, troubleshooting GPU usage, workload efficiency, cluster stability, and cost optimization.
  • Drive root cause analyses and corrective actions for complex cross-team production challenges.
  • Communicate insights to product and engineering teams to guide platform development and roadmap planning.
  • Create detailed documentation for architectures and implementation procedures and mentor team members as the group expands.

Required Skills and Experience

  • At least 8 years of experience in infrastructure, platform, or solutions engineering, including a minimum of 3 years specializing in AI/ML infrastructure or MLOps.
  • Expertise with Kubernetes including cluster lifecycle management, workloads, operator patterns, and RBAC policies.
  • Proficient with NVIDIA GPU hardware, ideally the latest generation accelerators.
  • Experience with distributed training technologies such as NCCL, tensor and pipeline parallelism, and LLM inference frameworks like vLLM and TGI.
  • Knowledge of GPU Operator, MIG, SR-IOV technologies, plus high-performance networking fabrics.
  • Strong scripting and automation capabilities, with preference for Python, Bash, and Go.
  • Hands-on familiarity with cloud platforms (AWS, Azure, GCP), including networking, identity and access management, and managed Kubernetes services.
  • Competence in observability tools like Prometheus, Grafana, and OpenTelemetry.
  • Excellent communication skills capable of interfacing effectively with engineering teams and executive stakeholders.
  • Additional advantages include experience with Run:AI or Slurm, expertise in GPU scheduling and autoscaling, or certifications such as CKA, CKAD, or cloud Solutions Architect credentials.

Tools & software

AWS Microsoft Azure Amazon Web Services AWS required Microsoft Azure required

How they work

Communication Problem Solving Leadership
🤖
Online · instant AI help
Broxer