Senior Site Reliability Engineer
Runware is building high-performance infrastructure and products to power the worlds intelligence. Our platform enables developers and businesses to run fast, scalable inference across image, video and emerging modalities, while our Serverless platform allows customers to deploy and scale their own AI models on production-grade GPU infrastructure.
As a Site Reliability Engineer at Runware, you will help ensure these systems remain reliable, performant and resilient as we scale. This is a highly technical, hands-on role working across software, infrastructure and production operations to improve observability, reduce incidents, eliminate operational toil and build lasting improvements across complex distributed systems.
What you’ll do
- Own and improve the reliability, availability and performance of critical production services across the Runware platform
- Define and evolve our reliability practices, including SLIs, SLOs, alerting, observability and production-readiness standards
- Investigate complex production issues across distributed systems, APIs, networking, queues, databases and GPU-backed workloads, participating in our engineering on-call rotation
- Lead and contribute to incident reviews and RCAs, turning recurring failure modes into lasting engineering improvements
- Reduce operational toil through automation, automated remediation and improvements to deployment safety, recovery and system resilience
- Work closely with Engineering and DevOps teams on capacity planning, performance, scaling and architectural improvements as the platform grows
- Have strong experience operating and troubleshooting production systems at scale in an SRE, Production Engineering, Platform Engineering or similar role
- Have a strong understanding of distributed systems and are comfortable debugging across applications, databases, queues, containers, networking and infrastructure
- Have experience designing and operating observability systems using metrics, logs and distributed tracing
- Understand SRE principles including SLIs, SLOs, error budgets, capacity planning, incident management and reducing operational toil
- Have experience with Kubernetes, containers, IaC and automated deployment practices, alongside the ability to write software and automation using languages such as Python, Go or PHP
- Take strong ownership of production problems and are comfortable participating in an engineering on-call rotation, taking issues from initial investigation through to long-term remediation
Bonus
- Experience operating high-throughput or low-latency APIs and distributed systems
- Experience with bare-metal infrastructure, GPU environments or AI and ML workloads
- Experience with RabbitMQ or other distributed messaging and queueing systems
- Experience operating MySQL, Redis, ClickHouse or similar production data systems
- Experience with global traffic management, load balancing, CDN platforms and hybrid infrastructure environments
- Experience building automated scaling, capacity management or self-healing systems
We’re a remote-first collective, meeting in person twice a year to plan, brainstorm, celebrate wins, and enjoy some face-to-face time. We have core hours for cooperative working and calls, but outside of that your calendar is yours. Work the hours that let you perform at your peak while also building a healthy life.
Our release cycles are fast and intense, but they’re followed by real downtime. After big pushes we expect the team to unplug, recharge, and come back ready & stronger than ever for the next leap.
- Generous paid time off – vacation, sick days, public holidays
- Meaningful stock options – share in the upside you create
- Remote-first setup – work from home anywhere we can employ you
- Flexible hours – own your schedule outside core collaboration blocks
- Family leave – paid maternity, paternity, and caregiver time
- Company retreats – twice-yearly gatherings in inspiring locations
Similar roles
Software Engineer II, Foundation
Chainalysis·London, Canada
Platform Engineering at Chainalysis provides the shared infrastructure, tooling, and developer experience layer that powers every engineering team. We are moving away from team-owned infrastructure silos and consolidating onto shared, scalable platform patterns. The Foundation discipline owns the substrate: cloud infrastructure, networking, Kubernetes, streaming, and security primitives. Everything else runs on what we build. We are hiring a Software Engineer…
- On-site
- Full-time
- Express Entry — PR day one, no employer
Software Engineer
Faculty·London, United Kingdom
WHY FACULTY? We established Faculty in 2014 because we thought that AI would be the most important technology of our time. Since then, we’ve worked with over 350 global customers to transform their performance through human-centric AI. You can read about our real-world impact here. We don’t chase hype cycles. We innovate, build and deploy responsible AI which moves the…
- On-site
- Full-time
Staff Analyst (f/m/d)
adjoe·Hamburg, Germany
adjoe builds the technologies behind mobile apps growth and monetization. With our core product Playtime Arcade, we've become the global leader in rewarded advertising, an ad unit built on a simple premise: users earn real in-app rewards for engaging with new apps. The result is one of the most effective value exchanges in adtech, connecting advertisers and publishers with over…
- On-site
- Full-time
- Blue Card — tied; settle 21–33mo
Senior Data Analyst - Experimentation & A/B Testing (f/m/d)
adjoe·Hamburg, Germany
adjoe builds the technologies behind mobile apps growth and monetization. With our core product Playtime Arcade, we've become the global leader in rewarded advertising, an ad unit built on a simple premise: users earn real in-app rewards for engaging with new apps. The result is one of the most effective value exchanges in adtech, connecting advertisers and publishers with over…
- On-site
- Full-time
- Blue Card — tied; settle 21–33mo
Data Operations & Labeling Specialist (all genders)
Stark·Munich, Germany
About Us STARK is a new kind of defence technology company revolutionizing the way autonomous systems are deployed across multiple domains. We design, develop and manufacture high-performance unmanned systems that are software-defined, mass-scalable, and cost-effective. This provides our operators with a decisive edge in highly contested environments. We're focused on delivering deployable, high-performance systems — not future promises. In a…
- On-site
- Full-time
- Blue Card — tied; settle 21–33mo
Software Developer | Full-stack (m/w/d)
Hoppe Marine Gmbh·Hamburg, Germany
Aufgaben Das Applications-Team entwickelt und prüft Lösungen, die hochwertige Sensordaten von großen Fracht- und Offshore-Schiffen erfassen – Daten, auf die sich unsere Kunden über SCADA, das Web und andere bordgestützte Anwendungen verlassen. Im Mittelpunkt unserer Arbeit steht das Software Engineering Center, unsere unternehmensinterne Kernanwendung, die unternehmensweit zur Konzeption und Konfiguration von Softwareprojekten für Schiffe auf der ganzen Welt eingesetzt wird.…
- On-site
- Full-time
- Blue Card — tied; settle 21–33mo
Werkstudent (m/w/d) Mechanische Konstruktion
Hoppe Marine Gmbh·Hamburg, Germany
Aufgaben • Du erstellst Fertigungszeichnungen und Stücklisten • Du fertigst technischer Dokumentationsunterlagen an • Du legst Artikel im unserem ERP-System an • Du führst Prototypenbau und Prototypentest durch Profil • Du bist Student:in im Bereich Maschinenbau, Mechatronik, Fahrzeugbau oder Schiffsbetriebstechnik • Du hast Erfahrung mit 3D-CAD Systemen (Autodesk Inventor wünschenswert) • Du arbeitest strukturiert, eigenständig und bist kommunikativ • Deine…
- On-site
- Full-time
- Blue Card — tied; settle 21–33mo
Security Engineer (all genders)
Xitaso·Augsburg, Germany
Werde Teil unserer Mission Unsere Kunden entwickeln komplexe Software- und Systemlösungen mit hohen Anforderungen an Cybersicherheit und Zuverlässigkeit. Dabei geht es nicht nur darum, Schwachstellen zu identifizieren, sondern Systeme so zu verstehen, dass nachhaltige Sicherheitslösungen entstehen. Du trägst dazu bei, indem du Systeme analysierst, Risiken bewertest und gemeinsam konkrete Verbesserungen umsetzt. Typische Aufgaben aus IT-Betrieb, Infrastruktur-Monitoring oder SOC stehen bei…
- On-site
- Full-time
- Blue Card — tied; settle 21–33mo