
Senior Software Engineer - Incident Insights & Readiness
The Incident Insights & Readiness SRE team at Datadog fosters a resilient culture by using incidents as learning opportunities and catalysts for growth. Our users are Datadog engineers, and we build the software, tooling, and operational frameworks that help them prepare for, respond to, and learn from incidents. We work closely with engineering teams across Datadog to analyze incidents and turn those insights into better tools, stronger incident response, and organizational learning. Our efforts empower Datadog to navigate unexpected failures confidently, efficiently, and with a commitment to continuous learning and systems improvement.
At Datadog, we place value in our office culture - the relationships and collaboration it builds and the creativity it brings to the table. We operate as a hybrid workplace to ensure our Datadogs can create a work-life harmony that best fits them.
The Incident Insights & Readiness SRE team at Datadog fosters a resilient culture by using incidents as learning opportunities and catalysts for growth. Our users are Datadog engineers, and we build the software, tooling, and operational frameworks that help them prepare for, respond to, and learn from incidents. We work closely with engineering teams across Datadog to analyze incidents and turn those insights into better tools, stronger incident response, and organizational learning. Our efforts empower Datadog to navigate unexpected failures confidently, efficiently, and with a commitment to continuous learning and systems improvement.
At Datadog, we place value in our office culture - the relationships and collaboration it builds and the creativity it brings to the table. We operate as a hybrid workplace to ensure our Datadogs can create a work-life harmony that best fits them.
What You’ll Do:
- Own and improve the on-call experience for the company by establishing best practices and building platforms to support on-call rotations and compensation.
- Define how we respond to incidents, lead the design and implementation of software to streamline the process, and collaborate with product teams to improve incident response across Datadog. Our aim is to fully support our incident responders in dealing with complexity.
- Contribute to the post-mortem process for the company, collaborating with teams on writing them, and identifying opportunities to reduce friction and enhance learning value for the organization. Our team also runs a weekly postmortem reading group.
- Support various teams in facilitating incident reviews that emphasize learning and blamelessness. Help them share their learnings across the organization to improve the resilience of our people.
- Provide technical leadership and day-to-day coaching to team members, accelerating their growth through design reviews, collaborative problem-solving and operational excellence best practices.
- Train our on-callers in incident and post-mortem processes, sharing expertise in incident management best practices. This involves both introducing newcomers to on-call responsibilities and refreshing the knowledge of existing engineers.
- Lead cross-functional initiatives in engineering organizations across Datadog, embedding with teams to understand their challenges and drive lasting improvements to reliability and operational excellence.
Who You Are:
- At least 5 years of experience building software that solves real user problems. Experience designing new features and collaborating on code and technical design reviews. We primarily develop in Go and Python, with a bit of TypeScript.
- Experience building or operating distributed systems, with familiarity with Kubernetes and an understanding of complex failure modes.
- Demonstrated ability to independently own ambiguous technical problems from design through delivery while balancing long-term engineering quality with pragmatic execution.
- Experience analyzing incidents, identifying systemic risks, and driving engineering improvements informed by operational learnings.
- Experience participating in on-call rotations and improving incident response processes. Experience serving as an incident commander or incident coordinator is a plus.
- Empathy, collaboration, and communication skills in English to cultivate strong relationships across various teams in the organization
- Experience mentoring engineers, driving cross-functional initiatives, and influencing technical direction without relying on organizational authority.
- We welcome candidates from a variety of backgrounds, including software engineering, site reliability engineering, production engineering, infrastructure, and other roles focused on building reliable systems or improving incident response.
Datadog values people from all walks of life. We understand not everyone will meet all the above qualifications on day one. That's okay. If you’re passionate about technology and want to grow your skills, we encourage you to apply.
Benefits and Growth:
- New hire stock equity (RSUs) and employee stock purchase plan (ESPP)
- Continuous professional development, product training, and career pathing
- Intradepartmental mentor and buddy program for in-house networking
- An inclusive company culture, ability to join our Community Guilds (Datadog employee resource groups)
- Access to Inclusion Talks, our internal panel discussions
- Free, global mental health benefits for employees and dependents age 6+
- Competitive global benefits
Benefits and Growth listed above may vary based on the country of your employment and the nature of your employment with Datadog.
#LI-Hybrid
About Datadog:
Datadog is the leading observability and security platform for the AI era, providing businesses with unified visibility across the technology stack to manage complexity at scale. It brings applications, infrastructure, data, models, and security into one place, using AI to detect and resolve issues before they impact customers. Trusted globally by Fortune 500 companies and high-growth AI leaders, Datadog enables businesses to move faster with clarity and confidence. Learn more about #DatadogLife on Instagram, LinkedIn, and Datadog Learning Center.
Equal Opportunity at Datadog:
Datadog is proud to offer equal employment opportunity to everyone regardless of race, color, ancestry, religion, sex, national origin, sexual orientation, age, citizenship, marital status, disability, gender identity, veteran status, and other characteristics protected by law. We also consider qualified applicants regardless of criminal histories, consistent with legal requirements. Here are our Candidate Legal Notices for your reference.
Datadog endeavors to make our Careers Page accessible to all users. If you would like to contact us regarding the accessibility of our website or need assistance completing the application process, please complete this form. This form is for accommodation requests only and cannot be used to inquire about the status of applications.
Privacy and AI Guidelines:
Any information you submit to Datadog as part of your application will be processed in accordance with Datadog’s Applicant and Candidate Privacy Notice. For information on our AI policy, please visit Interviewing at Datadog AI Guidelines.
Similar roles
Staff Software Engineer, Time and Scheduling
Gusto, Inc.·Worldwide
---------------------------------------- About Gusto At Gusto, we're on a mission to grow the small business economy. We handle the hard stuff — payroll, health insurance, 401(k)s, and HR — so owners can focus on their craft and their customers. With teams in Denver, San Francisco, and New York, we support more than 500,000 small businesses nationwide and are building a workplace…
- Remote
- Full-time
Senior Site Reliability Engineer (Performance and Scalability)
Digital Zone·United Arab Emirates
Your mission is to make DigitalZone able to scale. You will build the platform's capacity to absorb campaign-level traffic spikes, and you will give every engineering team the tools, standards, and practices to load- and failure test their own systems. This is an enablement role at its core: you raise the reliability bar across the org by building capability, not…
- Remote
- Full-time
- Green Visa — self-sponsored, no tie
Software Developer - Risk Technology
Squarepointcapital·London, United Kingdom
<p><strong>Position Overview:</strong></p> <p>Risk Technology is a global team that designs, builds and maintains Squarepoint’s trading risk platform, which is responsible for trade capture, position management, profit/loss computation, inventory/locate management and internal order routing. These critical systems need to be performant, resilient, and capable of timely processing of high volumes of trading data in both live and historical scenarios, requiring solutions…
- On-site
- Full-time
Software Developer - Data Pipelines (Python)
Squarepointcapital·London, United Kingdom
<p><strong>Position Overview:</strong></p> <p>We are seeking an experienced Python developer to join our Alpha Data team, responsible for delivering a vast quantity of data served to users worldwide. You will be a cornerstone of a growing Data team, becoming a technical subject matter expert and developing strong working relationships with quant researchers, traders, and fellow colleagues across our Technology organization.</p> <p>Alpha…
- On-site
- Full-time
Junior Software Developer - Front-end
Squarepointcapital·London, United Kingdom
<p><strong>Please only apply to the one job you feel best fits your skillset and experience. If our team feels you are better suited for another role, we will reach out about the alternate opportunity.</strong></p> <p><strong>Position Overview:</strong></p> <p><span class="ui-provider a b c d e f g h i j k l m n o p q r s t u v…
- On-site
- Full-time
Junior QA Engineer
AoFrio·New Zealand
Description Welcome to our World of Cold! At AoFrio, we are global leaders in providing IoT solutions to the food and beverage industry. Our innovative technology and dedicated team have positioned us at the forefront of our market. We are proud leaders in hardware-enabled Software as a Service (SaaS) for commercial refrigeration, with cutting-edge IoT solutions used by major brands…
- Full-time
- Skilled Migrant (Residence) — PR, points-based, no single employer
Software Engineer Specialist – Integration & AI
AIA Group·New Zealand
Your Role with Us As a Software Engineer Specialist – Integration in our Technology team, you’ll play a key role in designing and delivering integration solutions that keep our systems connected, secure, and running smoothly. You’ll collaborate with architects, engineers, and business stakeholders to build robust, scalable platforms that support both strategic initiatives and day-to-day operations. This is an exciting…
- Full-time
- Skilled Migrant (Residence) — PR, points-based, no single employer
(Intern) AI Automation Engineer
Tacto·Munich, Germany
YOUR IMPACT As part of Solution Engineering at Tacto, you'll bridge our powerful platform with measurable customer value through data expertise. Working directly with customers and leads, you'll understand their unique supply chain challenges and implement tailored technical solutions. By integrating customer procurement data and configuring the platform to match their processes, you'll drive adoption and showcase immediate ROI (even…
- On-site
- Internship
- Blue Card — tied; settle 21–33mo