FactFinder brand banner

Senior Site Reliability Engineer / SRE – Kubernetes & Hybrid Cloud (m/f/d)

FactFinderBerlin, GermanyPosted 2h agoonsite
via Arbeitnow

Introduction

At a glance

  • Location &workmodel:Berlin, hybrid
  • Tech stack:Kubernetes on our own servers, Harvester (KubeVirt), Argo CD/Flux, Prometheus/Grafana, Longhorn/Ceph
  • Team:A growing SRE team – you report to our CTPO for now and to the Team Lead SREwe'rehiring next; two system administrators in Pforzheim run the physical hardware
  • Process:Intro call · take-home task (~2h) · 90-min tech interview with our developers · leadership conversation · meet the team
  • Languages:Fluent Englishrequired; German is a plus, nota must

Why this role is special

Most SRE jobs today mean clicking around a managed cloud console. This one doesn't. We run our own hardware in Frankfurt and are building a modern private cloud platform on Kubernetes and Harvester – on-prem by default, with elastic burst into the public cloud and the option to go cloud-only later. You won't inherit a finished SRE practice: you'll help define it, side by side with our Berlin development teams – and you won't do it alone, a Team Lead SRE hire is coming next.

SRE here is an enabling discipline: you build what our developers need to ship reliably, while two system administrators in Pforzheim run the physical hardware. And the impact is direct – our product discovery technology powers more than 2,000 European online shops (Intersport, SPAR, Douglas and more), handling billions of shopper queries a year. When product discovery is slow or down, our customers lose revenue in real time.

Your first 90 days

You get to know both products, join the on-call rotation with a buddy, and own your first reliability topic – SLOs for one product, alerting that actually helps at 3 a.m., or automating away a piece of toil. By day 90 you've shipped visible improvements and know where you want to take the platform next.

Your mission

  • Define and own SLOs, SLIs and error budgets; drive data-informed reliability decisions
  • Lead incident response end-to-end: fast detection, clear communication, blameless postmortems – and reduce whole classes of incidents structurally, not case by case
  • Eliminatetoil through automation andGitOps; evolve our observability (metrics, logs, traces, alerting, runbooks) across two different stacks
  • Help build our custom Kubernetes operator (CRDs) that makes stateful search clusters declarative, self-healing and safely upgradable – and roll out the auto-scaling (HPA/VPA, KEDA, clusterautoscaler) today's architecture makes hard
  • Plan capacity,performanceand cost across on-premises and cloud – including the large-catalogue and peak-season loads our merchants care about – and use AI tools wherever they measurably speed up diagnosis and operations

Your profile

Must-haves:

  • Kubernetes in production – built, not just used:you'veset up andmaintainedclusters on your own servers (e.g.kubeadm, RKE2, k3s) and know cluster lifecycle and upgrades – managed-only experienceisn'tenough for this role
  • Lived SRE practice: SLOs, error budgets, incident management,on-call
  • Hands-on experience withGitOpsor comparable infrastructure/deployment automation– experience with Argo CD or Flux is a strong plus
  • Solid observability skills– metrics, logs, traces, alerting that people trust
  • A strong automation instinct–you'drather fix a problem's cause than repeat its workaround
  • A collaborative, enabling mindset– you see SRE as a service to our developers: you ask what they need, discuss trade-offs openly, anddon'tfall in love with your own solution

Nice-to-haves (genuinely optional – we'll teach you the rest):

  • Harvester,KubeVirt, vSphere/ESXi, OpenStack or similar virtualization/HCI platforms
  • Container storage (Longhorn, Ceph) and datacenter networking (load balancing, ingress, VLAN)
  • Auto-scaling (HPA, VPA, KEDA, clusterautoscaler) and capacity/cost planning
  • Experience building Kubernetes operators/CRDs
  • German language skills Certifications (CKA, CKS) are welcome but no substitute for hands-on experience – in the tech interview we'll ask about what you've actually built and operated.

You don't tick every box – or your title was never “SRE”? Apply anyway. If you've owned production systems, handled incidents and worked deeply with Kubernetes, we want to hear from you – production experience and engineering mindset matter more to us than titles or buzzwords.

THE JOY OF WORKING WITH US

  • Impact from day one: Your work directly influences the revenue of leading eCommerce brands across Europe.
  • Modern tech stack: Kubernetes, Harvester, GitOps, auto-scaling, and an exciting path toward the cloud – with room to build things right.
  • AI-first mindset: We use AI as a real part of our daily work, not as a buzzword.
  • Ownership & growth: Clear responsibility, short decision paths, and the opportunity to actively shape your role.
  • Flexible work: Hybrid work model three office days per week with a focus on outcomes.
  • Strong team: Experienced engineers, an open feedback culture, and an environment where reliability is treated as a real engineering discipline.

Job Location

Berlin, Munich, Pforzheim or Stockholm (all Hybrid)

Find Jobs in Germany on Arbeitnow

Skills

  • Software Development

Similar roles

  • HE

    Manager Housekeeping (m/w/d)

    HexenburgbeiDresden·Großharthau, Germany

    Sie haben ein Auge fürs Detail, lieben den Gästekontakt und behalten selbst an dynamischen Tagen den Überblick? Unser charmantes Ensemble aus 10 hochwertigen Ferienwohnungen, einer Reitanlage sowie exklusiven Eventbereichen sucht eine motivierte und organisierte Persönlichkeit, die unser Housekeeping-Team leitet und gleichzeitig als freundliches Gesicht an der (digitalen) Rezeption agiert. Anstellungsart: Vollzeit, Teilzeit Aufgaben 1. Leitung Housekeeping & Qualitätssicherung Teamführung &…

    • Full-time
    • Blue Cardtied; settle 21–33mo
  • Animore logo

    Senior Robotics Engineer (F/M/D)

    Animore·Munich, Germany

    THE OPPORTUNITY We're growing our robotics team and hiring for two closely related specializations: low-level control for mobile manipulators, and navigation/mobile manipulation. Both roles are deeply hands-on, close to real hardware, and central to getting our mobile manipulators from prototype toward general deployment. This is a role for someone who has spent their career close to actual robots in production.…

    • Full-time
    • Blue Cardtied; settle 21–33mo
  • Fiscal.ai logo

    Software Engineer: Quality (Terminal)

    Fiscal.ai·Toronto, Canada

    Job Title: Software Engineer: Quality (Terminal) Salary: $100,000-$220,000 + equity options. Location: Remote with occasional in-person work in our co-working space in downtown Toronto. About Us Fiscal.ai (formerly FinChat) is a leading research and data platform for capital markets. Combining a powerful research Terminal with modern APIs, Fiscal.ai is building the modern financial data company. The firm has raised $13M…

    • Remote
    • Full-time
    • Express EntryPR day one, no employer
  • Animore logo

    Working Student/ Intern - Physical AI & Robotic Systems Integration (F/M/D)

    Animore·Munich, Germany

    The Opportunity We are looking for a Physical AI & Robotic Systems Integration (Intern/Working Student) who thrives at the intersection of production software engineering, advanced AI, and physical hardware. If you find it incredibly satisfying to write clean, maintainable Python code that translates directly into precise, real-world robot movement, this role is for you. You will play a vital part…

    • Temp
    • Blue Cardtied; settle 21–33mo
  • Rockstardevelopers GmbH logo

    Machine Learning Engineer / MLOps Engineer (m/w/d)

    Rockstardevelopers GmbH·Stuttgart, Germany

    Du baust produktive KI-Systeme (LLM, RAG, Agenten) für große Kunden im öffentlichen Sektor. Kein Prototyp für die Schublade, sondern Software, die im Regelbetrieb läuft. Wer wir sind Rockstardevelopers, gegründet 2015, Büros in Stuttgart und München. Aus hunderten Projekten für Enterprise- und Mittelstandskunden wissen wir, wie man Software baut, die im Ernstfall trägt. Seit einer Weile verschiebt sich unser Schwerpunkt Richtung…

    • Remote
    • Full-time
    • Blue Cardtied; settle 21–33mo
  • 4lu44n1n37w012k logo

    Senior Indirect Procurement Category Manager - Marketing (All Genders)

    4lu44n1n37w012k·Berlin, Germany

    <h2><strong>The role</strong></h2> <p>The (Senior) Category Manager Marketing will be responsible for overseeing strategic sourcing and contracting at HelloFresh for Marketing across different marketing categories, which may include: Offline Media, Online Media, Influencer Marketing, Sponsorships, Retail Media, etc. The goal is to collaborate seamlessly with Global and regional stakeholders and major suppliers in all HelloFresh markets and help them to navigate…

    • On-site
    • Full-time
    • Blue Cardtied; settle 21–33mo
  • celonis logo

    Senior Application Product Manager - Supply Chain - Procurement

    celonis·Munich, Germany

    <div class="content-intro"><p>Celonis is the trusted platform to industrialize Enterprise AI. At our core is the Celonis Context Model — which combines process data, business knowledge, and intelligence into a living digital twin of the enterprise that AI can actually understand, turning AI's operational blind spots into operational clarity. World's leading companies trust Celonis and its global ecosystem of partners to…

    • On-site
    • Full-time
    • Blue Cardtied; settle 21–33mo
  • ILF Consulting Engineers Germany GmbH logo

    Werkstudent Proposal Management (m/w/d)

    ILF Consulting Engineers Germany GmbH·Germany

    Ein Team, in dem man sich wohlfühlt: Wir sind ein mittelständisches Familienunternehmen mit mehr als 500 Mitarbeitenden an verschiedenen Standorten in Deutschland. Als Ingenieurs- und Beratungsdienstleister sind wir weltweit in Projekten tätig, welche die Energiewende vorantreiben, den Schutz der Umwelt fördern und zur Verbesserung unser aller Lebensqualität beitragen. Die Position ist in unserer Abteilung Proposal Management angesiedelt, die aktuell 5…

    • On-site
    • Full-time
    • Blue Cardtied; settle 21–33mo