F

Senior Site Reliability Engineer (SRE) – Kubernetes & Hybrid Cloud

FactFinder

Berlin, Germany · Full Time

Be the first to apply

Experience
Any
Salary
Openings
1
Posted
1 week ago
Work mode
In office
Resume
Required to apply

Where you'll work

Sign in to tell us what does and doesn't work for you here — it sharpens every match we show you.

Job description

Introduction

Based in Berlin, this role offers a hybrid work environment within a growing Site Reliability Engineering (SRE) team responsible for managing Kubernetes clusters on private hardware using technologies like Harvester, Argo CD/Flux, Prometheus/Grafana, and Longhorn/Ceph. Reporting to the CTPO initially and soon to the incoming Team Lead SRE, this position supports a high-impact product discovery platform serving over 2,000 European online shops.

Your Mission

  • Define, maintain, and monitor Service Level Objectives (SLOs), Service Level Indicators (SLIs), and error budgets to guide data-driven reliability decisions.
  • Lead all incident management stages including quick detection, transparent communication, and blameless post-incident analysis with a focus on preventing recurring issues.
  • Automate repetitive tasks and improve observability solutions such as metrics, logs, traces, alerting, and runbooks across varied technology stacks.
  • Develop a bespoke Kubernetes operator (CRDs) that enables declarative, self-healing, and safe upgrades of stateful search clusters and implement autoscaling (HPA/VPA, KEDA, cluster autoscaler) where current architecture is limited.
  • Assess and plan capacity, performance, and costs across on-premise and cloud infrastructure, handling large catalogue and peak seasonal demands, while integrating AI tools to accelerate diagnostics and operational efficiency.

Your Profile

Essential qualifications:

  • Extensive hands-on experience building and managing Kubernetes clusters on personal or company servers (using kubeadm, RKE2, k3s) including cluster lifecycle and upgrades; solely managed cloud experience is insufficient.
  • Solid background in SRE practices encompassing SLOs, error budgets, incident management, and participating in on-call rotations.
  • Familiarity with GitOps or equivalent infrastructure automation approaches; experience with Argo CD or Flux is a significant advantage.
  • Strong skills in observability including metrics, logging, tracing, and creating trustworthy alerting systems.
  • A proactive approach to automation, prioritizing fixing root causes over applying temporary fixes.
  • A collaborative and service-oriented mindset focused on enabling development teams and fostering open dialogue on trade-offs.

Beneficial but optional skills:

  • Knowledge of Harvester, KubeVirt, vSphere/ESXi, OpenStack, or comparable virtualization and hyper-converged infrastructure technologies.
  • Experience with container storage solutions such as Longhorn or Ceph, and data center networking concepts like load balancing, ingress, and VLANs.
  • Understanding autoscaling mechanisms including HPA, VPA, KEDA, and cluster autoscaler, plus experience in capacity and cost planning.
  • Expertise in designing Kubernetes operators and custom resource definitions.
  • German language skills (not mandatory).

Professional certifications like CKA or CKS are valued but do not replace real-world experience. The recruitment process includes an introduction call, a take-home task, a technical interview, leadership discussions, and team meetings. Fluency in English is required; German is a plus but not a must.

The Joy of Working with Us

  • Your contributions impact the financial success of top European eCommerce brands from the start.
  • Work with a modern technological stack aligned with future cloud strategies and AI-driven operations.
  • Enjoy ownership of responsibilities, short decision-making processes, and room for personal growth.
  • Flexible hybrid work model requiring three office days per week with a strong emphasis on results.
  • A supportive and experienced engineering team committed to an engineering-driven reliability culture.

Job Locations

Hybrid work available across Berlin, Munich, Pforzheim, or Stockholm.

About the Company

We provide a leading AI-powered Product Discovery Platform for European eCommerce, supporting over 2,000 B2B and B2C clients and managing billions of shopper queries annually with a gross merchandise value exceeding €160 billion. Established for more than two decades and backed by private equity with strategic management, we operate offices in Pforzheim, Berlin, Munich, and Stockholm.

How they work

Communication Teamwork & Collaboration Problem Solving

Languages

Ericsson
🤖
Online · instant AI help
Broxer