XpertDirect

ML Deployment Engineer

XpertDirect

Munich, Bavaria, Germany · Full Time

Be the first to apply

Experience
3+ yrs
Salary
—
Openings
1
Posted
1 week ago
Work mode
In office
Resume
Required to apply

Where you'll work

Sign in to tell us what does and doesn't work for you here — it sharpens every match we show you.

Job description

About the Role

Our client, a fast-growing deep technology and AI company based in Munich, Germany, is seeking a Machine Learning (ML) Deployment Engineer to build the critical production systems that transition ML models from experimental stages into highly scalable and dependable real-world services. This role sits at the crossroads of ML Engineering, MLOps, and platform engineering, focused on creating robust tooling and infrastructure that ensures deployments are repeatable, well-monitored, and ready for production deployment.

Key Responsibilities

  • Design and implement production deployment pipelines specifically for machine learning models.
  • Manage the deployment and operation of model-serving workloads on Kubernetes clusters.
  • Create scalable inference services by utilizing KServe technology.
  • Package and containerize ML workloads using Docker containers for streamlined deployment.
  • Develop deployment automation tools and scripts primarily in Python.
  • Handle model version control, artifact management, and deployment workflows with MLflow.
  • Construct CI/CD pipelines that automate testing and releasing of ML services.
  • Deploy and operate workloads across cloud environments including AWS and/or Google Cloud Platform.
  • Design and implement strategies for rollout, rollback, and versioning of ML models in production.
  • Enhance system reliability, scalability, and observability for ML deployments.
  • Automate the workflow from model approval to production endpoint activation.
  • Work in close collaboration with ML Engineers to put new models into production without burdening them with infrastructure management.

Qualifications and Skills

  • At least 3 years of professional experience in MLOps, ML Engineering, ML Infrastructure, Platform Engineering, or related fields.
  • Proficient in Python programming for automation and tooling.
  • Experienced with Kubernetes container orchestration.
  • Knowledge of Docker for containerizing applications.
  • Familiarity with KServe or other comparable model-serving technologies.
  • Hands-on experience managing ML lifecycle using MLflow.
  • Proficiency in deploying ML workloads on cloud platforms such as AWS and/or GCP.
  • Strong understanding and practical experience with CI/CD pipelines and best practices.
  • Comprehensive understanding of production machine learning system challenges and requirements.

Additional Desirable Skills

  • Experience with Argo CD and GitOps methodologies.
  • Familiarity with Kubeflow orchestration platform.
  • Knowledge of NVIDIA Triton Inference Server and Ray Serve.
  • Experience with ML frameworks such as PyTorch or TensorFlow.
  • Expertise in monitoring tools like Prometheus and OpenTelemetry.
  • Infrastructure as code using Terraform.
  • Experience with advanced deployment strategies like canary or blue-green deployments.
  • Managing GPU-enabled inference workloads.
  • Capabilities in model monitoring and drift detection.
  • Handling real-time inference API operations.

Tools & software

Docker · 2 to 5 years required Kubernetes · 2 to 5 years required Mlflow · 2 to 5 years required Amazon Web Services AWS required Google Cloud Platform · 2 to 5 years required
🤖
Online · instant AI help
Broxer