FactFinder brand banner

Team Lead - Site Reliability Engineering (all genders)

FactFinderBerlin, GermanyPosted 2h agoonsite
via Arbeitnow

Introduction

FACT-Finder entwickelt Product-Discovery-Technologie für den eCommerce und ist mit den Produkten Next Generation und Infinity bei führenden Online-Shops in Europa im Einsatz. Aktuell modernisieren wir unser Hosting konsequent in Richtung Kubernetes auf Harvester – als on-prem Hybrid mit der Option, mittelfristig vollständig in die Cloud zu skalieren. Als Team Lead Site Reliability Engineering (all genders) verantwortest du Reliability, Skalierbarkeit und Kosten unserer Hosting-Umgebungen, treibst diese Transformation end-to-end voran und führst das Team, das sie umsetzt.

Deine Aufgaben

  • Du verantwortest die operative Gesundheit unseres Hostings über On-Premise (Frankfurt, Stockholm) und Cloud hinweg – Verfügbarkeit, Performance, Incident Management.
  • Du treibst die Modernisierung in Richtung Kubernetes auf Harvester aktiv voran: Cluster-Topologie, Storage (Longhorn), Networking (VLAN, Load Balancing, Ingress), Backup und Disaster Recovery.
  • Du baust eine produktionsreife k8s-Plattform auf: Lifecycle, Upgrades, RBAC, Secrets, GitOps (Argo CD / Flux), Observability und Policy-Guardrails.
  • Du gestaltest den NG Search Operator (Custom Kubernetes Operator) und löst Auto-Scaling (HPA, VPA, KEDA, Cluster Autoscaler) für die aktuelle Architektur.
  • Du definierst unser On-Prem-Hybrid-Modell konkret: welche Workloads laufen wo, wie burst’en wir in die Cloud, wie halten wir Latenz und Kosten im Griff – und hältst die Architektur portabel genug für einen späteren Cloud-Only-Schritt.
  • Du verantwortest Kapazitätsplanung und Hosting-Kosten und machst Kosten zu einem gezielt steuerbaren Hebel.
  • Du führst und entwickelst unser derzeit 4-köpfiges Hosting-Team fachlich und disziplinarisch, verantwortest Performance und prägst die technische Standards- und Ownership-Kultur.
  • Du machst AI zum festen Bestandteil unserer Operations: Diagnose, Automatisierung, Monitoring und Insight.

Dein Profil

  • Fundierter Hintergrund in Infrastructure oder Platform Engineering über On-Premise und Cloud hinweg.
  • Hands-on Tiefe mit Kubernetes in Produktion: Cluster-Lifecycle, Upgrades, Networking, Storage, RBAC, Observability, GitOps-Delivery.
  • Nachweislich starke Führungserfahrung, exzellente Kommunikation und Stakeholder-Management.
  • Idealerweise praktische Erfahrung mit Harvester oder vergleichbaren HCI-/Virtualisierungsplattformen (KubeVirt, vSphere/ESXi, OpenStack).
  • Erfahrung mit einer echten Migration von Bare Metal / klassischen VMs auf eine k8s-basierte Plattform – inklusive stateful Workloads, Storage-Migration, Cutover und Rollback.
  • Sicherer Umgang mit Kubernetes Operators (Custom Controllers / CRDs), idealerweise für stateful Systeme wie Search, Datenbanken oder Streaming.
  • Solides Verständnis von Auto-Scaling-Primitiven (HPA, VPA, Cluster Autoscaler, KEDA) und deren Zusammenspiel mit Kapazitätsplanung.
  • Erfahrung mit On-Prem-Hybrid-Architekturen und der Verantwortung für Reliability, Kapazität und Kosten produktiver Systeme.
  • Hands-on Fluency im Einsatz von AI-Tools im operativen Betrieb.
  • Sehr gute Englischkenntnisse; Deutsch von Vorteil.

THE JOY OF WORKING WITH US

  • Impact from day one: Deine Arbeit wirkt direkt auf die Umsätze führender eCommerce-Marken in Europa.
  • Führungsrolle mit Gestaltungsspielraum: Du führst ein eingespieltes Team und gestaltest unsere Plattform in einer entscheidenden Phase unserer Transformation.
  • Moderner Tech-Stack: Kubernetes, Harvester, GitOps, Auto-Scaling und ein spannender Weg in Richtung Cloud – mit Raum, Dinge neu und richtig zu bauen.
  • AI-first Mindset: Wir nutzen AI nicht als Buzzword, sondern als festen Bestandteil unserer täglichen Arbeit.
  • Ownership & Wachstum: Klare Verantwortung, kurze Entscheidungswege und die Möglichkeit, deine Rolle aktiv mitzugestalten.
  • Flexibles Arbeiten: Hybrides Arbeitsmodell mit Fokus auf Ergebnisse.
  • Starkes Team: Erfahrene Engineers, offene Feedback-Kultur und ein Umfeld, in dem Reliability als Engineering-Disziplin ernst genommen wird.
  • Attraktive Benefits: Wettbewerbsfähiges Gehalt, moderne Ausstattung, Weiterbildungsbudget und regelmäßige Team-Events.

Standort

​Berlin, München, Pforzheim oder Stockholm (hybrid)

Find more English Speaking Jobs in Germany on Arbeitnow

Skills

  • Software Development

Similar roles

  • Morningstar logo

    QA Automation Engineer

    Morningstar·Mumbai, India

    This job is with Morningstar, an inclusive employer and a member of myGwork – the largest global platform for the LGBTQ business community. Please do not contact the recruiter directly. The Quality Assurance Engineer will play a critical role in ensuring the accuracy, reliability, and readiness of the firm's modernized Investment Calculation Services. This position combines deep data quality analysis,…

    • Full-time
  • Reddit logo

    Engineering Manager, Notifications Platform

    Reddit·United States

    Reddit is a community of communities. It’s built on shared interests, passion, and trust, and is home to the most open and authentic conversations on the internet. Every day, Reddit users submit, vote, and comment on the topics they care most about. With 100,000+ active communities and approximately 130 million daily active unique visitors, Reddit is one of the internet’s…

    • Full-time
  • Mozn logo

    Engineering Manager

    Mozn·India

    About Mozn MOZN is a leading Enterprise AI company enabling organizations to make informed decisions in two critical domains: Financial Crime Prevention and Enterprise Knowledge Intelligence. We’re a diverse, collaborative team of innovators united by a shared purpose: to build AI that delivers tangible business value, builds trust, and empowers people and organizations with augmented intelligence. Our culture is built…

    • Remote
    • Full-time
  • hellofresh logo

    (Senior) Director UX, Operations

    hellofresh·Berlin, Germany

    <p>&nbsp;</p> <h2>About the role</h2> <p>HelloFresh is reorganising UX to work AI-native — practitioners either go deep as platform and skills experts, or broad as generalists who move fluidly across product, design and engineering to drive experience quality and innovation. This role leads UX for our Operations organisation: the part of the business that bridges technology with physical fulfilment, procurement, quality…

    • On-site
    • Full-time
    • Blue Cardtied; settle 21–33mo
  • hellofresh logo

    Security Engineer (ITSec) (m,f,x)

    hellofresh·Berlin, Germany

    <h2><strong>The role</strong></h2> <p>We’re looking for a new teammate to join us on the journey of keeping HelloFresh a trusted name - someone with a passion for security and appetite for new challenges. Security Engineers work in a variety of ways to constantly iterate and improve HelloFresh’s security posture.&nbsp;</p> <p>You will be the first port of call for responding to any…

    • On-site
    • Full-time
    • Blue Cardtied; settle 21–33mo
  • hellofresh logo

    Director UX Design

    hellofresh·Berlin, Germany

    <h1><strong>About HelloFresh</strong></h1> <p>HelloFresh Group reaches more than 6 million customers across 16 markets through a portfolio of 8 consumer brands — from meal kits and ready-to-eat meals to pet food and specialty products. HelloTech, our global technology division, builds the digital consumer experience across Berlin, Warsaw, and North America: the apps, the personalization, the discovery moments, and everything in between…

    • On-site
    • Full-time
    • Blue Cardtied; settle 21–33mo
  • Reflecta gGmbH logo

    Kaufmännische Geschäftsführung / COO (all genders), 80%

    Reflecta gGmbH·Berlin, Germany

    Reflecta demokratisiert den Zugang zu Netzwerk, Wissen und Fördermitteln. Seit 2010 bringen wir Menschen zusammen, die die Gesellschaft zukunftsfähig machen wollen: zunächst mit Offline-Formaten, seit 2019 digital. Heute ist Reflecta ein lernendes Ökosystem mit einer Community von über 10.000 Zukunftsgestalter:innen. Sein jüngstes Kapitel ist der Fördermittelkompass, die größte strukturierte Förderdatenbank im deutschsprachigen Raum, im Abo genutzt von Organisationen aus Zivilgesellschaft,…

    • Remote
    • Full-time
    • Blue Cardtied; settle 21–33mo
  • Tandem logo

    Product Manager

    Tandem·Berlin, Germany

    About Tandem Tandem is the leading global language learning community, with over 35 million members from all around the world practicing languages together. The concept of learning from each other is not only core to the Tandem community, it is also at the heart of how we work every day at Tandem. We are based in Berlin, profitable and growing,…

    • On-site
    • Temp
    • Blue Cardtied; settle 21–33mo