Datadog Logo

Datadog

Senior Software Engineer - Incident Insights & Readiness

Posted 49 Minutes Ago
Be an Early Applicant
Easy Apply
Hybrid
New York, NY, USA
192K-240K Annually
Senior level
Easy Apply
Hybrid
New York, NY, USA
192K-240K Annually
Senior level
Build and improve software, platforms, and operational frameworks for incident response, on-call readiness, post-mortem learning, and engineering resilience. Lead incident-process design, facilitate reviews, mentor engineers, train on-callers, and drive cross-functional reliability initiatives. The role requires experience with distributed systems, Kubernetes, incident analysis, on-call operations, and software development primarily in Go and Python, with some TypeScript.
The summary above was generated by AI

We're on a mission to build the best platform in the world for engineers to understand and scale their systems, applications, and teams. We operate at high scale—trillions of data points per day—providing always-on alerting, metrics visualization, logs, and application tracing for tens of thousands of companies. Our engineering culture values pragmatism, honesty, and simplicity to solve hard problems the right way

 

The Incident Insights & Readiness SRE team at Datadog fosters a resilient culture by using incidents as learning opportunities and catalysts for growth. Our users are Datadog engineers, and we build the software, tooling, and operational frameworks that help them prepare for, respond to, and learn from incidents. We work closely with engineering teams across Datadog to analyze incidents and turn those insights into better tools, stronger incident response, and organizational learning. Our efforts empower Datadog to navigate unexpected failures confidently, efficiently, and with a commitment to continuous learning and systems improvement.

 

At Datadog, we place value in our office culture - the relationships and collaboration it builds and the creativity it brings to the table. We operate as a hybrid workplace to ensure our Datadogs can create a work-life harmony that best fits them.

 

What You’ll Do: 

  • Own and improve the on-call experience for the company by establishing best practices and building platforms to support on-call rotations and compensation.

  • Define how we respond to incidents, lead the design and implementation of software to streamline the process, and collaborate with product teams to improve incident response across Datadog. Our aim is to fully support our incident responders in dealing with complexity.

  • Contribute to the post-mortem process for the company, collaborating with teams on writing them, and identifying opportunities to reduce friction and enhance learning value for the organization. Our team also runs a weekly postmortem reading group.

  • Support various teams in facilitating incident reviews that emphasize learning and blamelessness. Help them share their learnings across the organization to improve the resilience of our people.

  • Provide technical leadership and day-to-day coaching to team members, accelerating their growth through design reviews, collaborative problem-solving and operational excellence best practices.

  • Train our on-callers in incident and post-mortem processes, sharing expertise in incident management best practices. This involves both introducing newcomers to on-call responsibilities and refreshing the knowledge of existing engineers.

  • Lead cross-functional initiatives in engineering organizations across Datadog, embedding with teams to understand their challenges and drive lasting improvements to reliability and operational excellence.

 

Who You Are: 

  • At least 5 years of experience building software that solves real user problems. Experience designing new features and collaborating on code and technical design reviews. We primarily develop in Go and Python, with a bit of TypeScript.

  • Experience building or operating distributed systems, with familiarity with Kubernetes and an understanding of complex failure modes.

  • Demonstrated ability to independently own ambiguous technical problems from design through delivery while balancing long-term engineering quality with pragmatic execution.

  • Experience analyzing incidents, identifying systemic risks, and driving engineering improvements informed by operational learnings.

  • Experience participating in on-call rotations and improving incident response processes. Experience serving as an incident commander or incident coordinator is a plus.

  • Empathy, collaboration, and communication skills in English to cultivate strong relationships across various teams in the organization

  • Experience mentoring engineers, driving cross-functional initiatives, and influencing technical direction without relying on organizational authority.

  • We welcome candidates from a variety of backgrounds, including software engineering, site reliability engineering, production engineering, infrastructure, and other roles focused on building reliable systems or improving incident response.

 

Datadog values people from all walks of life. We understand not everyone will meet all the above qualifications on day one. That's okay. If you’re passionate about technology and want to grow your skills, we encourage you to apply.

 

Benefits and Growth: 

  • New hire stock equity (RSUs) and employee stock purchase plan (ESPP)

  • Continuous professional development, product training, and career pathing

  • Intradepartmental mentor and buddy program for in-house networking

  • An inclusive company culture, ability to join our Community Guilds (Datadog employee resource groups)

  • Access to Inclusion Talks, our internal panel discussions

  • Free, global mental health benefits for employees and dependents age 6+

  • Competitive global benefits

 

Benefits and Growth listed above may vary based on the country of your employment and the nature of your employment with Datadog.

Datadog offers a competitive salary and equity package, and may include variable compensation. Actual compensation is based on factors such as the candidate's skills, qualifications, and experience. In addition, Datadog offers a wide range of best in class, comprehensive and inclusive employee benefits for this role including healthcare, dental, parental planning, and mental health benefits, a 401(k) plan and match, paid time off, fitness reimbursements, and a discounted employee stock purchase plan.

The reasonably estimated yearly salary for this role at Datadog is:
$192,000—$240,000 USD

About Datadog: 

Datadog is the leading observability and security platform for the AI era, providing businesses with unified visibility across the technology stack to manage complexity at scale. It brings applications, infrastructure, data, models, and security into one place, using AI to detect and resolve issues before they impact customers. Trusted globally by Fortune 500 companies and high-growth AI leaders, Datadog enables businesses to move faster with clarity and confidence. Learn more about #DatadogLife on Instagram, LinkedIn, and Datadog Learning Center.

Equal Opportunity at Datadog:

Datadog is proud to offer equal employment opportunity to everyone regardless of race, color, ancestry, religion, sex, national origin, sexual orientation, age, citizenship, marital status, disability, gender identity, veteran status, and other characteristics protected by law. We also consider qualified applicants regardless of criminal histories, consistent with legal requirements. Here are our Candidate Legal Notices for your reference. 

Datadog endeavors to make our Careers Page accessible to all users. If you would like to contact us regarding the accessibility of our website or need assistance completing the application process, please complete this form. This form is for accommodation requests only and cannot be used to inquire about the status of applications. 

Privacy and AI Guidelines:

Any information you submit to Datadog as part of your application will be processed in accordance with Datadog’s Applicant and Candidate Privacy Notice. For information on our AI policy, please visit Interviewing at Datadog AI Guidelines.

HQ

Datadog New York, New York, USA Office

We are located in the New York Times building and five-minute walk away from Times Square. The 42 St Port Authority Bus Terminal is right across the street, providing a highly accessible transportation network.

Similar Jobs at Datadog

2 Hours Ago
Easy Apply
Hybrid
New York, NY, USA
Easy Apply
120K-160K Annually
Senior level
120K-160K Annually
Senior level
Artificial Intelligence • Cloud • Security • Software • Cybersecurity
Own full-cycle technical recruiting, including sourcing, screening, interview coordination, candidate follow-up, compensation negotiation, and closing. Partner with leadership to improve interview processes, train interview teams, maintain accurate Greenhouse data, and deliver an excellent candidate experience. The role requires recruiting software engineering talent, managing multiple openings independently, and operating effectively in a fast-changing environment.
Top Skills: Greenhouse
4 Hours Ago
Easy Apply
Hybrid
New York, NY, USA
Easy Apply
112K-140K Annually
Mid level
112K-140K Annually
Mid level
Artificial Intelligence • Cloud • Security • Software • Cybersecurity
Optimize Datadog’s technical written and video content for traditional search, answer engines, generative AI, and agentic platforms. Responsibilities include on-page SEO, metadata, internal linking, keyword research, content gap analysis, content briefs, AI visibility analysis, cross-functional collaboration, and performance reporting. The role requires technical content experience, hands-on SEO expertise, familiarity with AEO/GEO, and proficiency with SEO tools and AI-assisted search platforms.
Top Skills: AeoAgentic PlatformsAhrefsAi-Assisted SearchChatgptClaudeCodexConductorGeminiGeoGoogle Search ConsoleProfoundSemrushSeo
Yesterday
Easy Apply
Hybrid
New York, NY, USA
Easy Apply
131K-164K Annually
Senior level
131K-164K Annually
Senior level
Artificial Intelligence • Cloud • Security • Software • Cybersecurity
Build and facilitate global recruiting and people enablement programs across the recruiting lifecycle. Responsibilities include creating curricula, facilitator guides, participant materials, job aids, asynchronous modules, knowledge checks, and LMS/SCORM packages; delivering live training; measuring completion and readiness; incorporating AI into manager development; and leaving a maintainable content system for Recruiting Operations.
Top Skills: AILmsScorm

What you need to know about the NYC Tech Scene

As the undisputed financial capital of the world, New York City is an epicenter of startup funding activity. The city has a thriving fintech scene and is a major player in verticals ranging from AI to biotech, cybersecurity and digital media. It also has universities like NYU, Columbia and Cornell Tech attracting students and researchers from across the globe, providing the ecosystem with a constant influx of world-class talent. And its East Coast location and three international airports make it a perfect spot for European companies establishing a foothold in the United States.

Key Facts About NYC Tech

  • Number of Tech Workers: 549,200; 6% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Capgemini, Bloomberg, IBM, Spotify
  • Key Industries: Artificial intelligence, Fintech
  • Funding Landscape: $25.5 billion in venture capital funding in 2024 (Pitchbook)
  • Notable Investors: Greycroft, Thrive Capital, Union Square Ventures, FirstMark Capital, Tiger Global Management, Tribeca Venture Partners, Insight Partners, Two Sigma Ventures
  • Research Centers and Universities: Columbia University, New York University, Fordham University, CUNY, AI Now Institute, Flatiron Institute, C.N. Yang Institute for Theoretical Physics, NASA Space Radiation Laboratory

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account