AI Inference Engineer (all genders)
Karlsruhe, Baden-Württemberg, Germany · Full Time
Be the first to apply
- Experience
- Any
- Salary
- —
- Openings
- 1
- Posted
- 4 weeks ago
- Work mode
- In office
- Resume
- Required to apply
Where you'll work
Sign in to tell us what does and doesn't work for you here — it sharpens every match we show you.
Job description
Position Overview
Join as an AI Inference Engineer to build the foundational technology for operational AI systems within regulated environments. You will design, develop, and manage LLM inference platforms that operate on-premises or in private cloud setups, ensuring security, scalability, observability, and cost efficiency.
Key Responsibilities
- Design, develop and maintain productive LLM inference platforms meeting strict demands for data sovereignty, security, and operational control, deployed on-premises, in private clouds, or sovereign European cloud environments.
- Collaborate with cloud, platform, security, and data engineering teams together with customers to transition AI use-cases into full production.
- Integrate contemporary inference engines and open-weight models within Kubernetes, container, and platform infrastructures.
- Plan and optimize GPU and memory resources alongside inference workloads considering model size, quantization, batching, KV-cache strategies, latency, throughput, and cost-efficiency.
- Oversee runtime environments of live AI systems, including model serving, APIs, authentication, secrets management, observability, and logging.
- Extract reusable reference architectures, deployment templates, and operational playbooks from client projects to advance the company's applied AI capabilities.
Candidate Profile
- Professional background with experience in platform engineering, cloud infrastructure, MLOps, LLMOps, DevOps, backend engineering, or machine learning engineering.
- Deep understanding of technical and economic aspects related to modern LLM inference, including GPU utilization, model serving, quantization, batching, KV-cache management, latency, throughput, and cost control.
- Proficiency in container and cloud tools and platforms such as Docker, Kubernetes, Helm, Terraform, CI/CD pipelines, Linux, and monitoring solutions.
- Strong understanding of AI concepts, especially Transformer-based models like large language models and embeddings, enabling informed decisions for practical AI system deployment.
- Keen awareness of security and governance topics including identity management, permissions, secrets, logging, auditing, and compliance, especially in regulated sectors.
- Effective communication skills capable of explaining complex technology clearly; pragmatic approach and confident navigation in dynamic project environments.
- Additional advantages include experience with vLLM, SGLang, similar inference technologies, GPU clusters, sovereign cloud environments, or private clouds.
- Willingness to travel flexibly across Germany on customer sites.
- Fluency in German and English is required to thrive in this role and company culture.
Why Join Exxeta
Exxeta is a place where digital solutions are developed to make a real impact in companies, markets, and minds. Our team of over 1200 professionals combines technology, ideas, and diverse perspectives fueled by curiosity, team spirit, and a strong drive for meaningful results. We value diversity, different viewpoints, attitude, ideas, and hands-on motivation. Join us to be part of a vibrant and impactful environment.