Senior Site Reliability Engineer (SRE) – Kubernetes & Hybrid Cloud
Berlin, Germany · Full Time
Be the first to apply
- Experience
- Any
- Salary
- —
- Openings
- 1
- Posted
- 1 week ago
- Work mode
- In office
- Resume
- Required to apply
Where you'll work
Sign in to tell us what does and doesn't work for you here — it sharpens every match we show you.
Job description
Introduction
Based in Berlin, this role offers a hybrid work environment within a growing Site Reliability Engineering (SRE) team responsible for managing Kubernetes clusters on private hardware using technologies like Harvester, Argo CD/Flux, Prometheus/Grafana, and Longhorn/Ceph. Reporting to the CTPO initially and soon to the incoming Team Lead SRE, this position supports a high-impact product discovery platform serving over 2,000 European online shops.
Your Mission
- Define, maintain, and monitor Service Level Objectives (SLOs), Service Level Indicators (SLIs), and error budgets to guide data-driven reliability decisions.
- Lead all incident management stages including quick detection, transparent communication, and blameless post-incident analysis with a focus on preventing recurring issues.
- Automate repetitive tasks and improve observability solutions such as metrics, logs, traces, alerting, and runbooks across varied technology stacks.
- Develop a bespoke Kubernetes operator (CRDs) that enables declarative, self-healing, and safe upgrades of stateful search clusters and implement autoscaling (HPA/VPA, KEDA, cluster autoscaler) where current architecture is limited.
- Assess and plan capacity, performance, and costs across on-premise and cloud infrastructure, handling large catalogue and peak seasonal demands, while integrating AI tools to accelerate diagnostics and operational efficiency.
Your Profile
Essential qualifications:
- Extensive hands-on experience building and managing Kubernetes clusters on personal or company servers (using kubeadm, RKE2, k3s) including cluster lifecycle and upgrades; solely managed cloud experience is insufficient.
- Solid background in SRE practices encompassing SLOs, error budgets, incident management, and participating in on-call rotations.
- Familiarity with GitOps or equivalent infrastructure automation approaches; experience with Argo CD or Flux is a significant advantage.
- Strong skills in observability including metrics, logging, tracing, and creating trustworthy alerting systems.
- A proactive approach to automation, prioritizing fixing root causes over applying temporary fixes.
- A collaborative and service-oriented mindset focused on enabling development teams and fostering open dialogue on trade-offs.
Beneficial but optional skills:
- Knowledge of Harvester, KubeVirt, vSphere/ESXi, OpenStack, or comparable virtualization and hyper-converged infrastructure technologies.
- Experience with container storage solutions such as Longhorn or Ceph, and data center networking concepts like load balancing, ingress, and VLANs.
- Understanding autoscaling mechanisms including HPA, VPA, KEDA, and cluster autoscaler, plus experience in capacity and cost planning.
- Expertise in designing Kubernetes operators and custom resource definitions.
- German language skills (not mandatory).
Professional certifications like CKA or CKS are valued but do not replace real-world experience. The recruitment process includes an introduction call, a take-home task, a technical interview, leadership discussions, and team meetings. Fluency in English is required; German is a plus but not a must.
The Joy of Working with Us
- Your contributions impact the financial success of top European eCommerce brands from the start.
- Work with a modern technological stack aligned with future cloud strategies and AI-driven operations.
- Enjoy ownership of responsibilities, short decision-making processes, and room for personal growth.
- Flexible hybrid work model requiring three office days per week with a strong emphasis on results.
- A supportive and experienced engineering team committed to an engineering-driven reliability culture.
Job Locations
Hybrid work available across Berlin, Munich, Pforzheim, or Stockholm.
About the Company
We provide a leading AI-powered Product Discovery Platform for European eCommerce, supporting over 2,000 B2B and B2C clients and managing billions of shopper queries annually with a gross merchandise value exceeding €160 billion. Established for more than two decades and backed by private equity with strategic management, we operate offices in Pforzheim, Berlin, Munich, and Stockholm.