E

AI Inference Engineer

Exxeta

Karlsruhe, Baden-Württemberg, Germany · Full Time

Be the first to apply

Experience
Any
Salary
—
Openings
1
Posted
4 days ago
Work mode
In office
Resume
Required to apply

Where you'll work

Sign in to tell us what does and doesn't work for you here — it sharpens every match we show you.

Job description

Role Overview

As an AI Inference Engineer, you will establish the technical foundation for operational AI systems within regulated settings. You will be responsible for creating and maintaining large language model (LLM) inference platforms deployed either on-premises or within private cloud infrastructures, ensuring they operate securely, scalably, with observability, and cost efficiency.

Key Responsibilities

  • Design, develop, and operate high-demand LLM inference platforms that emphasize data sovereignty, security, and operational control, whether deployed on-premises, in private clouds, or sovereign European cloud environments.
  • Collaborate closely with cloud, platform, security, and data engineering teams alongside clients to transition AI use cases into reliable production environments.
  • Integrate modern inference engines and open-weight models within Kubernetes, containerized, and various platform environments.
  • Plan and optimize GPU and memory resources along with inference workloads, balancing factors such as model size, quantization, batching, KV-cache strategies, latency, throughput, and cost.
  • Manage the runtime of production AI systems including model serving, APIs, authentication, secrets management, observability, and logging capabilities.
  • Create reusable reference architectures, deployment templates, and operational playbooks derived from client projects to enhance the Applied AI competencies of the team.

Candidate Profile

  • Background: Demonstrated experience in platform engineering, cloud infrastructure, MLOps, LLMOps, DevOps, backend engineering, or machine learning engineering. Proven ability to build and operate production-level systems and a strong drive for rapid personal development.
  • Inference Engineering Expertise: Deep understanding of technical and economic aspects of modern LLM inference including model serving, GPU utilization, quantization, batching, KV-cache management, latency, throughput, and cost factors.
  • Cloud and Platform Skills: Proficiency in Docker, Kubernetes, Helm, Terraform, CI/CD pipelines, Linux environments, and system observability tools.
  • AI Knowledge: Solid comprehension of transformer-based models like LLMs and embeddings, enabling well-founded technical decisions for AI production systems.
  • Security and Governance Awareness: Consideration of identity management, authorizations, secrets handling, logging, auditing, and compliance—especially within regulated environments—from the outset.
  • Communication and Work Style: Ability to clearly explain complex technical topics, pragmatic mindset, and confidence navigating dynamic project settings.
  • Additional Assets: Experience with vLLM, SGLang, similar inference technologies, GPU clusters, and sovereign or private cloud environments is a plus.
  • Mobility: Willingness to travel and provide onsite consulting nationwide.
  • Languages: Fluency in both German and English is required to thrive within the company culture and operations.

Why Join Exxeta

At Exxeta, we create digital solutions that profoundly impact businesses, markets, and mindsets. Our over 1200 colleagues bring together technology, ideas, and diverse perspectives driven by curiosity, teamwork, and a commitment to make meaningful impact. We foster an environment rich in diversity and varying viewpoints where mindset, ideas, and enthusiasm to take action are paramount.

Tools & software

Docker required Kubernetes required Terraform required

How they work

Communication Teamwork & Collaboration Problem Solving Adaptability Customer Focus

Languages

Servicenow Ericsson

Leave it if you'd like a reply — we won't use it for anything else.

Click to browse, drag & drop, or paste a screenshot

PNG, JPG, GIF, MP4, WebM, MOV · Max 20MB each · Up to 5 files

🤖
Online · instant AI help
Broxer