X

Senior MLOps Engineer (AI Platform)

Xcede

Berlin, Germany (Hybrid) · Full Time

Be the first to apply

Experience
4+ yrs
Salary
EUR 85,000 – EUR 110,000 / year
Openings
1
Posted
5 days ago
Work mode
Hybrid
Resume
Required to apply

Where you'll work

Sign in to tell us what does and doesn't work for you here — it sharpens every match we show you.

Job description

About the Company

The client is a prominent global energy and sustainability company operating in 20 countries, employing over 6,000 staff and generating revenue exceeding €1 billion. Their solutions serve 14 million residential and commercial properties worldwide.

Role Overview

As part of the AI platform engineering team, you will be responsible for maintaining and advancing the infrastructure that supports the company's AI agents and large language model (LLM) applications. Your primary focus will be on deployment, system reliability, observability, and cost efficiency, ensuring that both managed and self-hosted AI models operate securely, consistently, and efficiently in live production environments. You will define the operational framework and standards upon which the broader AI engineering group relies.

Key Responsibilities

  • Develop and manage deployment infrastructure tailored for LLM systems, handling both managed and self-hosted models.
  • Administer cloud infrastructure using infrastructure-as-code approaches, particularly with Terraform.
  • Implement containerization and orchestration frameworks, mainly leveraging Kubernetes.
  • Lead efforts in observability, including monitoring, logging, and detection of model and data drift within production environments.
  • Provide expert advice on model selection, optimize performance, and manage operational costs.
  • Define and enforce production standards encompassing version control, rollback strategies, compliance, and security.
  • Collaborate closely with AI engineers to transition prototypes and models into reliable, scalable production systems.

Required Qualifications and Experience

  • Extensive senior-level experience with MLOps or specifically LLMOps, operating production ML and LLM systems.
  • Proficiency in infrastructure-as-code management, especially with Terraform.
  • Strong background with Kubernetes and container orchestration technologies.
  • Cloud platform experience, preferably with Microsoft Azure, though AWS or Google Cloud Platform experience is also applicable.
  • Solid expertise in Python programming alongside sound software engineering principles such as continuous integration and deployment, testing, and code quality assurance.
  • A production-focused approach emphasizing monitoring, cost management, governance, and system reliability.
  • Exposure to Java is advantageous but not mandatory.
  • Language proficiency: minimum C1 level German and strong English communication skills.

Compensation and Benefits

  • Annual salary ranging from €85,000 to €110,000.
  • 30 days of paid annual leave.
  • Hybrid working model permitting up to 50% remote work.
  • Flexible working hours to support work-life balance.
  • Culture that prioritizes continuous learning with ongoing training and development opportunities.
  • Agile working environment within a well-established international company.

Tools & software

Kubernetes · 2 to 5 years required Terraform · 2 to 5 years required

How they work

Teamwork & Collaboration Work Ethic Learning Agility Results Orientation

Languages

Servicenow Ericsson
🤖
Online · instant AI help
Broxer