XpertDirect

AI Inference Platform Engineer

XpertDirect

Berlin, Germany · Full Time

Be the first to apply

Experience
4+ yrs
Salary
Openings
1
Posted
1 week ago
Work mode
In office
Resume
Required to apply

Where you'll work

Sign in to tell us what does and doesn't work for you here — it sharpens every match we show you.

Job description

Role Overview

Our growing AI Infrastructure client based in Berlin is seeking an AI Inference Platform Engineer to enhance and maintain the platform powering the deployment of production AI models on GPU-accelerated systems. The role focuses on the interplay of AI infrastructure, distributed systems, and platform engineering to maximize inference efficiency, GPU capacity, scaling, response times, and platform robustness.

Key Responsibilities

  • Build and manage Kubernetes clusters to support AI model inference in production environments.
  • Deploy and fine-tune model-serving workloads using vLLM and NVIDIA Triton Inference Server.
  • Enhance GPU resource usage, increase throughput, and minimize latency during inference operations.
  • Create and implement autoscaling solutions tailored to dynamic AI task demands.
  • Develop automation and tooling primarily using Python to streamline platform processes.
  • Configure and oversee infrastructure provisioning leveraging Terraform.
  • Establish comprehensive observability across AI models, GPU hardware, Kubernetes environments, and inference services.
  • Diagnose and resolve bottlenecks affecting inference performance.
  • Upgrade batching, concurrency, caching, and resource management approaches for optimal performance.
  • Design dependable deployment procedures for launching new models and updating versions.
  • Collaborate closely with machine learning engineers to facilitate smooth transition of models into production systems.

Required Qualifications

  • Minimum 4 years experience in AI Infrastructure, ML Infrastructure, MLOps, platform engineering, or equivalent roles.
  • Proficiency with Kubernetes cluster management and orchestration.
  • Strong programming skills in Python for building tooling and automation.
  • Hands-on experience with NVIDIA GPU hardware and associated ecosystem.
  • Expertise deploying vLLM and NVIDIA Triton Inference Server for model serving.
  • Familiarity with infrastructure provisioning using Terraform.
  • Knowledge of observability tools and monitoring systems.
  • Solid understanding of Linux operating systems and distributed production environments.

Desirable Skills

  • Experience with CUDA programming and NVIDIA GPU Operator deployments.
  • Working knowledge of PyTorch, KServe, and Ray Serve frameworks.
  • Skills with monitoring platforms such as Prometheus, Grafana, and OpenTelemetry.
  • Expertise in optimizing large language model (LLM) inference including quantization techniques.
  • Multi-GPU inference strategies and management.
  • Familiarity with cloud GPU infrastructure on AWS or GCP.
  • Track record operating inference services that demand high throughput or low latency.

Tools & software

Leave it if you'd like a reply — we won't use it for anything else.

Click to browse, drag & drop, or paste a screenshot

PNG, JPG, GIF, MP4, WebM, MOV · Max 20MB each · Up to 5 files

🤖
Online · instant AI help
Broxer