
Staff Site Reliability Engineer
Who We Are
AI is changing how software gets built. Code production is becoming a commodity. The focus is shifting from writing code to orchestrating, verifying, and governing change – and the toolchain is the new constraint.
We are at the center of this shift. We build Develocity, a toolchain observability and intelligence platform used by some of the world's leading software organizations – Netflix, Airbnb, Spotify, SAP, major global banks, and hundreds more. Develocity helps software teams achieve delivery excellence through deep observability, build and test acceleration, and AI-powered intelligence across the entire toolchain – with current support for Gradle Build Tool, Apache Maven™, sbt, npm, and Python.
We are an AI-native company. AI is not a feature we're bolting on – it's central to how we work, how we think about our product, and where we're heading. We're investing deeply in making Develocity's unique data and decades of domain expertise accessible to both humans and AI agents, with trust, evidence, and explainability at the core of everything we build.
We have partnered with the Apache Software Foundation, the Commonhaus Foundation, the Micronaut Foundation, and other OSS projects such as Spring, Quarkus, Kotlin, JUnit, AndroidX, and many more to bring the values of Develocity also to the OSS Community.
Our Values
Seek to Understand: Everything starts with listening and understanding; we strive to understand diverse viewpoints, problems, and motivations. Before we take action, we ensure we truly grasp the challenges, perspectives, and goals.
Know the Why: We approach our work with a clear sense of purpose, ensuring every step is deliberate and focused. We take meaningful action with urgency, but never at the expense of thoughtful consideration.
Innovate & Iterate: We embrace challenges and are not afraid to try new things, even if they might fail. With a deep understanding and a clear purpose, we can develop creative, bold solutions to tackle challenges.
Own the Outcome: We are empowered to take initiative, and we maintain transparency in our work and its outcomes. When we execute, we take responsibility for our decisions, measure the success of our innovations, and learn from the results.
Who You Are
We're building a new SRE team and looking for founding members to help shape how we operate. As a Lead SRE, you'll be a technical and operational leader for reliability across Develocity. You'll help define our SRE vision, set standards for how we operate production services, and mentor other SREs as the team grows. This is a hands-on role with broad influence across engineering, cloud platform, and customer-facing teams.
The SRE team will be responsible for the reliability, performance, and availability of Develocity instances serving paying customers, open-source projects, and public-facing services, plus supporting infrastructure like artifact registries.
You'll work on our internally-built Cloud Application Platform, Kubernetes on AWS, and develop deep expertise in it. When incidents happen, you'll troubleshoot issues across the stack, from application to infrastructure. You'll collaborate with the Cloud Platform team to improve the tooling you depend on, and with engineering teams to build reliability into how we ship software. If you like automating things and hate doing the same task twice, you'll fit in well.
You'll be part of a distributed, remote-first team that values asynchronous communication and written documentation. Strong self-direction and clear communication across time zones are essential.
Responsibilities
- Operate and maintain all Develocity instances and supporting services in production.
- Define and evolve SRE standards, practices, and operating models, including on-call, incident response, postmortems, and SLOs.
- Participate in a follow-the-sun on-call rotation, acting as a technical escalation point for complex or high-severity incidents.
- Lead incident response and blameless retrospectives, ensuring learnings result in measurable reliability improvements.
- Set reliability priorities using risk, customer impact, business goals, SLOs, and error budgets.
- Identify systemic reliability risks and continuously evolve Develocity's SaaS operations as the platform and customer base grow.
- Lead and influence architectural and design reviews to ensure reliability, scalability, and operability.
- Drive automation across deployment, upgrades, monitoring, self-healing, recovery, and operational workflows.
- Build and maintain comprehensive observability for all managed services, including logging, metrics, tracing, and alerting.
- Own disaster recovery, backups, and business continuity planning and execution.
- Partner with engineering leadership to balance feature delivery with reliability and operational excellence.
- Mentor and coach SREs, supporting technical growth and strong operational practices.
- Help onboard new SREs and contribute to hiring by defining and assessing SRE excellence at Develocity.
- Communicate clearly with customers during incidents and maintenance windows.
- Optimize performance, resource utilization, and operational costs.
Minimum qualifications
- 7+ years in SRE, DevOps, or an equivalent role operating production services at scale.
- Experience leading reliability initiatives across multiple teams or services.
- Demonstrated ability to influence technical direction without direct authority.
- Experience designing and operating systems with SLOs and error budgets, and exercising strong judgment in balancing reliability, velocity, and cost.
- Strong Kubernetes experience in production environments.
- Cloud infrastructure expertise, preferably AWS (EKS, RDS, S3, EC2).
- Proficiency with observability tools (Prometheus, Grafana) and Infrastructure as Code (Terraform).
- Track record of incident management and response in a 24/7 on-call environment.
- Scripting proficiency (Python, Bash) for automation.
- Strong written and verbal English communication skills.
Preferred qualifications
- Experience as a founding or early SRE establishing practices in a growing SaaS organization.
- Familiarity with Develocity.
- JVM language experience (Java, Kotlin).
- Experience with customer-facing and executive-level incident communications.
What We Offer
- A ground-floor role in a new SRE team - you'll shape how we do things, not inherit someone else's decisions.
- Real ownership of production systems used by engineers at companies you've heard of.
- Direct interaction with customers when things go wrong (and when they go right).
- A culture that values automation over heroics.
- In-person meetings, such as our annual company offsite and team meetings.
- Work from home in a remote-first environment.
- Competitive salaries and equity grants.
Location
- Remote from anywhere in Europe (GMT).
- While our team works remotely and is spread across the globe, we deeply value daily interactions and collaboration.
Find Jobs in United Kingdom on Arbeitnow
Similar roles
Lead Consultant - Civil Nuclear (Deputy Practice Lead)
Decision Analysis Services Ltd·Bristol, United Kingdom
· Location: Manchester or Bristol, with travel to client sites · Contract: Full Time, Permanent *Please be aware that all offers of employment will be subject to a UK Security Clearance check. To gain this you ordinarily need at least 6 years’ UK residency. * Overview Decision Analysis Services Limited (DAS) is an independent professional services company. Since 2007 we…
- Hybrid
- Full-time
Lead Consultant - Civil Nuclear (Deputy Practice Lead)
Decision Analysis Services Ltd·Manchester, United Kingdom
· Location: Manchester or Bristol, with travel to client sites · Contract: Full Time, Permanent *Please be aware that all offers of employment will be subject to a UK Security Clearance check. To gain this you ordinarily need at least 6 years’ UK residency. * Overview Decision Analysis Services Limited (DAS) is an independent professional services company. Since 2007 we…
- Hybrid
- Full-time
Warehouse Operative
N2O·Bedford, United Kingdom
We are looking to recruit warehouse operatives to join our growing team in Bedford. The focus of the role is to assist the Warehouse Manager and Team Leaders with the day to day running of the warehouse, to ensure the safe storage, inventory, pick & packing and delivery deadlines are met, and the warehouse is a safe environment to work…
- Full-time
Cleaner/Driver
ABM UK·Horley, United Kingdom
JOB TITLE: Aircraft Cleaner LOCATION: London Gatwick Airport RH6 0NP CONTRACT: Permanent, 40 hours per week, 4 on 4 off PAY RATE: Competitive ROLE OVERVIEW AND PURPOSE The main purposes of the Aircraft Cleaner role are to work efficiently, as directed by the Cleaning Supervisors, as part of a team you will assist cleaning and restock the aircraft to the…
- Part-time
People Advisor - 12 month FTC
Doctor Care Anywhere·London, United Kingdom
Thanks for stopping by! We’re Doctor Care Anywhere (DCA): The UK’s largest private provider of telehealth services. The Company works with insurers, healthcare providers and corporate customers to serve patients with a range of digitally enabled telehealth services on its proprietary platform. DCA is committed to delivering the best possible patient experience and clinical care through digitally enabled, joined up,…
- Hybrid
- Full-time
Support Worker (Glasgow East)
HRM Homecare Services·Glasgow, United Kingdom
HRM Homecare Services is seeking compassionate and dedicated Support Workers to join our team in Glasgow East The Support Worker role involves providing essential care and support to individuals within their own homes, promoting independence, dignity, and quality of life. Key Responsibilities: * Provide personal care and assistance with daily living activities such as bathing, dressing, eating, and medication support.…
- Contract
Multiskilled Operative
ABM UK·Edinburgh, United Kingdom
LOCATION: Gyle Shopping Centre SHIFT PATTERN: 5 days over 7, 40 hours per week PAY RATE: Competitive If you require any additional support or adjustments during the recruitment process, please don't hesitate to contact our Recruitment Department at [email protected]. We're here to help! As part of the ABM Service Team at Gyle Shopping Centre, you will be responsible for delivering…
- Full-time
Event Content Producer - Ecosystems
Founders Forum Group·London, United Kingdom
This role sits at the heart of our Ecosystems events portfolio, bringing to life innovative, community-focused events that push beyond our traditional forum model. You'll work across our most experimental and niche-focused events - from emerging new concepts in early exploration through to established community-driven programmes - managing research, stakeholder relationships, content curation, and operational execution. If you're looking to…
- Full-time