Focused brand banner

Site Reliability Engineer II - AI & Infrastructure (f/m/d)

FocusedBerlin, GermanyPosted 2h agoonsite
via Arbeitnow

Focused Energy is a pioneering international deep-tech company with locations in Germany and the US, dedicated to commercializing laser-driven nuclear fusion. Our mission is to deliver clean, safe, and virtually limitless energy to the world. As we rapidly scale, we seek talented individuals who thrive on bringing clarity and structure to fast-growing environments.

About the Role

Focused Energy is looking for a Site Reliability Engineer II to help build reliable, secure, observable, and easy-to-operate systems for its internal applications. You’ll own day-to-day reliability, deployments, incident response, monitoring, backup, recovery, and rollback processes across a broad technical environment spanning cloud services, databases, identity, networking, secrets, and CI/CD.

 

This role is ideal for an engineer who enjoys solving complex operational problems and turning recurring incidents into lasting improvements. Alongside your core infrastructure responsibilities, you’ll provide overflow support for AI-enablement workflows and tools when the dedicated AI Tech Enabler needs additional support.

 

What You’ll Do

  • Own the reliability and day-to-day operation of multiple internal applications, deployment platforms, and supporting services.

  • Lead the investigation and resolution of incidents and deployment failures across cloud services, databases, identity, networking, secrets, and CI/CD systems.

  • Drive automation improvements across build, release, deployment, monitoring, alerting, backup, recovery, rollback, and runbook processes.

  • Implement safe, well-tested code, configuration, infrastructure, and database fixes to restore or improve service reliability.

  • Develop proposals for cloud, hosting, database, and runtime migrations, as well as horizontal scaling approaches, supported by appropriate testing and rollback plans.

  • Build and maintain clear operational documentation, including runbooks, recovery procedures, incident fixes, and system knowledge that can be reused by other engineers.

  • Partner with software engineering, IT, security, identity, and external platform-support stakeholders to plan and execute infrastructure changes safely.

  • Provide overflow diagnostic and troubleshooting support for AI workflows and tools, including Langdock, without losing focus on core reliability priorities.

 

Who You Are

  • You take ownership of well-scoped technical work from investigation through testing, deployment, and follow-up.

  • You are calm and methodical when responding to live incidents and can diagnose problems across multiple interconnected systems.

  • You focus on root-cause resolution rather than repeatedly applying temporary fixes.

  • You communicate incidents, trade-offs, risks, and recovery plans clearly to both technical and non-technical stakeholders.

  • You are proactive about identifying reliability risks, automation opportunities, recurring issues, and operational improvements.

  • You make independent decisions on tactical fixes and safe configuration changes while seeking appropriate sign-off for architecture, migration, and scaling decisions.

  • You document changes and operational knowledge so that other engineers can support systems effectively.

  • You ask for help early when risks or dependencies are unclear and keep the Head of IT informed about significant issues and progress.

 

Desirable Skills & Knowledge

  • Strong experience with Linux, networking, structured troubleshooting, and cloud hosting concepts.

  • Experience operating internal applications and deployment platforms in Azure or a comparable cloud environment.

  • Practical knowledge of relational databases, particularly PostgreSQL or an equivalent platform, including safe migration and recovery practices.

  • Experience with infrastructure as code and CI/CD tools such as OpenTofu, Terraform, GitLab CI, GitHub Actions, or Azure DevOps.

  • Knowledge of monitoring, alerting, backup, recovery, rollback automation, incident response, and service reliability practices.

  • Experience with horizontal scaling, stateless system design, migrations, and reversible change management.

  • Scripting and automation skills, together with experience reducing repetitive operational work.

  • Familiarity with identity and access management, least-privilege access, secrets management, FE’s internal AI/SaaS stack, and tools such as Langdock.


Focused Energy is an equal opportunity employer committed to creating an inclusive environment. Qualified applicants will receive consideration for employment without regard to race, color, religion, sex, sexual orientation, gender perception or identity, national origin, age, marital status, protected veteran status, or disability status.

Pursuant to the San Francisco Fair Chance Ordinance, Focused Energy will consider for employment qualified applicants with arrest and conviction records.

Compensation offered will be determined by factors such as location, level, job-related knowledge, skills, and experience. Certain roles may be eligible for incentive compensation, equity, benefits. 

Find more English Speaking Jobs in Germany on Arbeitnow

Skills

  • IT

Similar roles

  • es recruitment pte. ltd. logo

    Senior SRE Automation Engineer — Selenium/Playwright

    es recruitment pte. ltd.·Singapore, Singapore

    es recruitment pte. ltd. is seeking an Automation & Site Reliability Engineer (3 openings) for a banking technology team in Singapore. You will design and maintain automation solutions using Selenium/Playwright with Python/Java/TypeScript, build automation frameworks and ELK dashboards, and apply SRE practices to improve reliability. The role focuses on automation development, monitoring/observability and production support, with on-site work at Changi…

    • Contract
    • SGD 552–SGD 912 / yr
    • Employment Passtied to employer
  • Taxfix logo

    Senior Platform Engineer (d/f/m)

    Taxfix·Berlin, Germany

    OUR STORY: Every year millions of people are either filing their taxes in fear or giving up on their tax refund altogether. We're working on fixing that. Our intuitive app enables anyone, regardless of education or background, to file their taxes with newfound confidence. Spread across Germany, Spain and the UK, the team at Taxfix Group with its brands Taxfix,…

    • On-site
    • Full-time
    • Blue Cardtied; settle 21–33mo
  • ElevenLabs logo

    Customer Success - EMEA

    ElevenLabs·Germany

    ABOUT ELEVENLABS ElevenLabs is an AI research and product company transforming how we interact with technology. We launched in January 2023 with the first human-like AI voice model. Today, we serve millions of users and thousands of businesses - from fast-growing startups to large enterprises like Deutsche Telekom and Meta. Our investors are some of the world's most prominent, including…

    • On-site
    • Full-time
    • Blue Cardtied; settle 21–33mo
  • Moss logo

    Global Talent Acquisition Partner (f/m/d)

    Moss·Berlin, Germany

    At Moss, we give finance professionals the power to automate their day-to-day and make forward-thinking decisions. Our team and culture make us unique — we’re driven by impact and growth, where every one of us strives to learn and excel. Recognised by Sifted’s Rising 100 and LinkedIn's Top Startups, we’re here to help propel your career and together, make Moss…

    • On-site
    • Full-time
    • Blue Cardtied; settle 21–33mo
  • Statista logo

    Presales Manager DACH (m/f/d)

    Statista·Hamburg, Germany

    Bei Statista dreht sich alles um Daten und Fakten, denn wir sind die weltweit führende Business Data Plattform. Durch die Bereitstellung verlässlicher und einfach nutzbarer Daten sowie verschiedener Datenanalyseprodukte und -dienstleistungen unterstützen wir Menschen weltweit, faktenbasierte Entscheidungen zu treffen. 2007 in Hamburg gegründet, haben wir uns schnell zu einem globalen Unternehmen mit Büros in Metropolen wie London, New York und…

    • On-site
    • Full-time
    • Blue Cardtied; settle 21–33mo
  • Reddit logo

    Backend Software Engineer, PDP Experience

    Reddit·Worldwide

    Reddit is a community of communities. It’s built on shared interests, passion, and trust, and is home to the most open and authentic conversations on the internet. Every day, Reddit users submit, vote, and comment on the topics they care most about. With 100,000+ active communities and approximately 130 million daily active unique visitors, Reddit is one of the internet’s…

    • Remote
    • Full-time
  • CURRENTA GRUPPE logo

    Teamleiter Backoffice (m/w/d)

    CURRENTA GRUPPE·Leverkusen, Germany

    LERN UNS KENNEN Die Chemion Logistik GmbH ist ein Logistikdienstleister, der auf die Bedürfnisse der Chemie und chemienahen Geschäftszweige spezialisiert ist. Wir beschäftigen an den drei CHEMPARK-Standorten (Leverkusen, Dormagen und Krefeld-Uerdingen) insgesamt ca. 900 Mitarbeiter*innen in den Bereichen Transport, Lagerung, Umschlag, Expedition, Equipment und Schulung. #Traumjob, #TeamChemion, #neueChancen UNTERSTÜTZE UNS ALS TEAMLEITER BACKOFFICE (M/W/D) Wir suchen ab sofort einen Teamleiter…

    • On-site
    • Full-time
    • Blue Cardtied; settle 21–33mo
  • Jobgether logo

    Brand Protection Specialist (SEO)

    Jobgether·Germany

    This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Brand Protection Specialist (SEO) based in Germany. This is an opportunity to play a key role in protecting a global digital brand and its intellectual property across search engines and online channels. You will monitor the web…

    • On-site
    • Full-time
    • Blue Cardtied; settle 21–33mo