AI Inference Engineer
Karlsruhe, Baden-Württemberg, Germany · Full Time
Be the first to apply
- Experience
- Any
- Salary
- —
- Openings
- 1
- Posted
- 4 days ago
- Work mode
- In office
- Resume
- Required to apply
Where you'll work
Sign in to tell us what does and doesn't work for you here — it sharpens every match we show you.
Job description
Role Overview
As an AI Inference Engineer, you will establish the technical foundation for operational AI systems within regulated settings. You will be responsible for creating and maintaining large language model (LLM) inference platforms deployed either on-premises or within private cloud infrastructures, ensuring they operate securely, scalably, with observability, and cost efficiency.
Key Responsibilities
- Design, develop, and operate high-demand LLM inference platforms that emphasize data sovereignty, security, and operational control, whether deployed on-premises, in private clouds, or sovereign European cloud environments.
- Collaborate closely with cloud, platform, security, and data engineering teams alongside clients to transition AI use cases into reliable production environments.
- Integrate modern inference engines and open-weight models within Kubernetes, containerized, and various platform environments.
- Plan and optimize GPU and memory resources along with inference workloads, balancing factors such as model size, quantization, batching, KV-cache strategies, latency, throughput, and cost.
- Manage the runtime of production AI systems including model serving, APIs, authentication, secrets management, observability, and logging capabilities.
- Create reusable reference architectures, deployment templates, and operational playbooks derived from client projects to enhance the Applied AI competencies of the team.
Candidate Profile
- Background: Demonstrated experience in platform engineering, cloud infrastructure, MLOps, LLMOps, DevOps, backend engineering, or machine learning engineering. Proven ability to build and operate production-level systems and a strong drive for rapid personal development.
- Inference Engineering Expertise: Deep understanding of technical and economic aspects of modern LLM inference including model serving, GPU utilization, quantization, batching, KV-cache management, latency, throughput, and cost factors.
- Cloud and Platform Skills: Proficiency in Docker, Kubernetes, Helm, Terraform, CI/CD pipelines, Linux environments, and system observability tools.
- AI Knowledge: Solid comprehension of transformer-based models like LLMs and embeddings, enabling well-founded technical decisions for AI production systems.
- Security and Governance Awareness: Consideration of identity management, authorizations, secrets handling, logging, auditing, and compliance—especially within regulated environments—from the outset.
- Communication and Work Style: Ability to clearly explain complex technical topics, pragmatic mindset, and confidence navigating dynamic project settings.
- Additional Assets: Experience with vLLM, SGLang, similar inference technologies, GPU clusters, and sovereign or private cloud environments is a plus.
- Mobility: Willingness to travel and provide onsite consulting nationwide.
- Languages: Fluency in both German and English is required to thrive within the company culture and operations.
Why Join Exxeta
At Exxeta, we create digital solutions that profoundly impact businesses, markets, and mindsets. Our over 1200 colleagues bring together technology, ideas, and diverse perspectives driven by curiosity, teamwork, and a commitment to make meaningful impact. We foster an environment rich in diversity and varying viewpoints where mindset, ideas, and enthusiasm to take action are paramount.
Industry
IT Services & Consulting