
Senior Site Reliability Engineer / SRE – Kubernetes & Hybrid Cloud (m/f/d)
Introduction
At a glance
- Location &workmodel:Berlin, hybrid
- Tech stack:Kubernetes on our own servers, Harvester (KubeVirt), Argo CD/Flux, Prometheus/Grafana, Longhorn/Ceph
- Team:A growing SRE team – you report to our CTPO for now and to the Team Lead SREwe'rehiring next; two system administrators in Pforzheim run the physical hardware
- Process:Intro call · take-home task (~2h) · 90-min tech interview with our developers · leadership conversation · meet the team
- Languages:Fluent Englishrequired; German is a plus, nota must
Why this role is special
Most SRE jobs today mean clicking around a managed cloud console. This one doesn't. We run our own hardware in Frankfurt and are building a modern private cloud platform on Kubernetes and Harvester – on-prem by default, with elastic burst into the public cloud and the option to go cloud-only later. You won't inherit a finished SRE practice: you'll help define it, side by side with our Berlin development teams – and you won't do it alone, a Team Lead SRE hire is coming next.
SRE here is an enabling discipline: you build what our developers need to ship reliably, while two system administrators in Pforzheim run the physical hardware. And the impact is direct – our product discovery technology powers more than 2,000 European online shops (Intersport, SPAR, Douglas and more), handling billions of shopper queries a year. When product discovery is slow or down, our customers lose revenue in real time.
Your first 90 days
You get to know both products, join the on-call rotation with a buddy, and own your first reliability topic – SLOs for one product, alerting that actually helps at 3 a.m., or automating away a piece of toil. By day 90 you've shipped visible improvements and know where you want to take the platform next.
Your mission
- Define and own SLOs, SLIs and error budgets; drive data-informed reliability decisions
- Lead incident response end-to-end: fast detection, clear communication, blameless postmortems – and reduce whole classes of incidents structurally, not case by case
- Eliminatetoil through automation andGitOps; evolve our observability (metrics, logs, traces, alerting, runbooks) across two different stacks
- Help build our custom Kubernetes operator (CRDs) that makes stateful search clusters declarative, self-healing and safely upgradable – and roll out the auto-scaling (HPA/VPA, KEDA, clusterautoscaler) today's architecture makes hard
- Plan capacity,performanceand cost across on-premises and cloud – including the large-catalogue and peak-season loads our merchants care about – and use AI tools wherever they measurably speed up diagnosis and operations
Your profile
Must-haves:
- Kubernetes in production – built, not just used:you'veset up andmaintainedclusters on your own servers (e.g.kubeadm, RKE2, k3s) and know cluster lifecycle and upgrades – managed-only experienceisn'tenough for this role
- Lived SRE practice: SLOs, error budgets, incident management,on-call
- Hands-on experience withGitOpsor comparable infrastructure/deployment automation– experience with Argo CD or Flux is a strong plus
- Solid observability skills– metrics, logs, traces, alerting that people trust
- A strong automation instinct–you'drather fix a problem's cause than repeat its workaround
- A collaborative, enabling mindset– you see SRE as a service to our developers: you ask what they need, discuss trade-offs openly, anddon'tfall in love with your own solution
Nice-to-haves (genuinely optional – we'll teach you the rest):
- Harvester,KubeVirt, vSphere/ESXi, OpenStack or similar virtualization/HCI platforms
- Container storage (Longhorn, Ceph) and datacenter networking (load balancing, ingress, VLAN)
- Auto-scaling (HPA, VPA, KEDA, clusterautoscaler) and capacity/cost planning
- Experience building Kubernetes operators/CRDs
- German language skills Certifications (CKA, CKS) are welcome but no substitute for hands-on experience – in the tech interview we'll ask about what you've actually built and operated.
You don't tick every box – or your title was never “SRE”? Apply anyway. If you've owned production systems, handled incidents and worked deeply with Kubernetes, we want to hear from you – production experience and engineering mindset matter more to us than titles or buzzwords.
THE JOY OF WORKING WITH US
- Impact from day one: Your work directly influences the revenue of leading eCommerce brands across Europe.
- Modern tech stack: Kubernetes, Harvester, GitOps, auto-scaling, and an exciting path toward the cloud – with room to build things right.
- AI-first mindset: We use AI as a real part of our daily work, not as a buzzword.
- Ownership & growth: Clear responsibility, short decision paths, and the opportunity to actively shape your role.
- Flexible work: Hybrid work model three office days per week with a focus on outcomes.
- Strong team: Experienced engineers, an open feedback culture, and an environment where reliability is treated as a real engineering discipline.
Job Location
Berlin, Munich, Pforzheim or Stockholm (all Hybrid)
Find Jobs in Germany on Arbeitnow
Skills
- Software Development
Similar roles
Manager Housekeeping (m/w/d)
HexenburgbeiDresden·Großharthau, Germany
Sie haben ein Auge fürs Detail, lieben den Gästekontakt und behalten selbst an dynamischen Tagen den Überblick? Unser charmantes Ensemble aus 10 hochwertigen Ferienwohnungen, einer Reitanlage sowie exklusiven Eventbereichen sucht eine motivierte und organisierte Persönlichkeit, die unser Housekeeping-Team leitet und gleichzeitig als freundliches Gesicht an der (digitalen) Rezeption agiert. Anstellungsart: Vollzeit, Teilzeit Aufgaben 1. Leitung Housekeeping & Qualitätssicherung Teamführung &…
- Full-time
- Blue Card — tied; settle 21–33mo
Senior Robotics Engineer (F/M/D)
Animore·Munich, Germany
THE OPPORTUNITY We're growing our robotics team and hiring for two closely related specializations: low-level control for mobile manipulators, and navigation/mobile manipulation. Both roles are deeply hands-on, close to real hardware, and central to getting our mobile manipulators from prototype toward general deployment. This is a role for someone who has spent their career close to actual robots in production.…
- Full-time
- Blue Card — tied; settle 21–33mo
Software Engineer: Quality (Terminal)
Fiscal.ai·Toronto, Canada
Job Title: Software Engineer: Quality (Terminal) Salary: $100,000-$220,000 + equity options. Location: Remote with occasional in-person work in our co-working space in downtown Toronto. About Us Fiscal.ai (formerly FinChat) is a leading research and data platform for capital markets. Combining a powerful research Terminal with modern APIs, Fiscal.ai is building the modern financial data company. The firm has raised $13M…
- Remote
- Full-time
- Express Entry — PR day one, no employer
Working Student/ Intern - Physical AI & Robotic Systems Integration (F/M/D)
Animore·Munich, Germany
The Opportunity We are looking for a Physical AI & Robotic Systems Integration (Intern/Working Student) who thrives at the intersection of production software engineering, advanced AI, and physical hardware. If you find it incredibly satisfying to write clean, maintainable Python code that translates directly into precise, real-world robot movement, this role is for you. You will play a vital part…
- Temp
- Blue Card — tied; settle 21–33mo
Machine Learning Engineer / MLOps Engineer (m/w/d)
Rockstardevelopers GmbH·Stuttgart, Germany
Du baust produktive KI-Systeme (LLM, RAG, Agenten) für große Kunden im öffentlichen Sektor. Kein Prototyp für die Schublade, sondern Software, die im Regelbetrieb läuft. Wer wir sind Rockstardevelopers, gegründet 2015, Büros in Stuttgart und München. Aus hunderten Projekten für Enterprise- und Mittelstandskunden wissen wir, wie man Software baut, die im Ernstfall trägt. Seit einer Weile verschiebt sich unser Schwerpunkt Richtung…
- Remote
- Full-time
- Blue Card — tied; settle 21–33mo
Senior Indirect Procurement Category Manager - Marketing (All Genders)
4lu44n1n37w012k·Berlin, Germany
<h2><strong>The role</strong></h2> <p>The (Senior) Category Manager Marketing will be responsible for overseeing strategic sourcing and contracting at HelloFresh for Marketing across different marketing categories, which may include: Offline Media, Online Media, Influencer Marketing, Sponsorships, Retail Media, etc. The goal is to collaborate seamlessly with Global and regional stakeholders and major suppliers in all HelloFresh markets and help them to navigate…
- On-site
- Full-time
- Blue Card — tied; settle 21–33mo
Senior Application Product Manager - Supply Chain - Procurement
celonis·Munich, Germany
<div class="content-intro"><p>Celonis is the trusted platform to industrialize Enterprise AI. At our core is the Celonis Context Model — which combines process data, business knowledge, and intelligence into a living digital twin of the enterprise that AI can actually understand, turning AI's operational blind spots into operational clarity. World's leading companies trust Celonis and its global ecosystem of partners to…
- On-site
- Full-time
- Blue Card — tied; settle 21–33mo
Werkstudent Proposal Management (m/w/d)
ILF Consulting Engineers Germany GmbH·Germany
Ein Team, in dem man sich wohlfühlt: Wir sind ein mittelständisches Familienunternehmen mit mehr als 500 Mitarbeitenden an verschiedenen Standorten in Deutschland. Als Ingenieurs- und Beratungsdienstleister sind wir weltweit in Projekten tätig, welche die Energiewende vorantreiben, den Schutz der Umwelt fördern und zur Verbesserung unser aller Lebensqualität beitragen. Die Position ist in unserer Abteilung Proposal Management angesiedelt, die aktuell 5…
- On-site
- Full-time
- Blue Card — tied; settle 21–33mo