Kontakt.io Logo

Kontakt.io

Lead Site Reliability Engineer

Posted 6 Hours Ago
Be an Early Applicant
Hybrid
New York, NY, USA
200K-250K Annually
Expert/Leader
Hybrid
New York, NY, USA
200K-250K Annually
Expert/Leader
Lead reliability, performance, automation, and scalability for a 24/7 AWS-based healthcare platform. Responsibilities include designing fault-tolerant infrastructure, managing Kubernetes and Docker environments, implementing Terraform and GitOps, improving observability, leading incident response and disaster recovery, and ensuring HIPAA and SOC 2 compliance. The role also sets technical strategy, reduces downtime through self-healing automation, and leads and mentors the SRE team.
The summary above was generated by AI

About Kontakt.io

Every day, intelligent software orchestrates the physical world around us – from matching drivers with riders to optimizing global supply chains. Yet inside hospitals, where every second and every decision can affect a patient's life, operations are still spread across dozens of disconnected systems.

At Kontakt.io, we're changing that. By combining proprietary hardware, AI-powered intelligence, and deep integrations with the systems hospitals already rely on, we're creating a real-time understanding of hospital operations that software alone can't deliver. That intelligence powers the execution layer hospitals have been missing – helping care teams make smarter decisions and deliver better patient care.

Backed by Goldman Sachs and trusted by leading health systems including HCA Healthcare, Sutter Health, AdventHealth, Trinity Health, Northwell Health, Cleveland Clinic, and the U.S. Department of Veterans Affairs, we're pioneering the next generation of healthcare operations. We've more than doubled our revenue over the past year and are on track to surpass $70M in annual recurring revenue – not because we're following a market, but because we're defining one.

If you're excited to solve hard problems, work with a team of builders, and help hospitals deliver better care every day, we'd love to meet you!

We’re looking for a Lead Site Reliability Engineer to own the reliability, performance, and automation of our cloud-based, real-time platform. This role will focus on keeping our platform running smoothly 24/7, minimizing downtime, improving observability, incident response, and self-healing automation. You will lead and scale the SRE team to ensure our infrastructure stays ahead of demand, operates efficiently, and meets the needs of our growing healthcare customers.

What You'll Do

  • Ensure 99.99% uptime across our cloud platform, meeting strict SLAs for healthcare customers.
  • Design and implement self-healing, fault-tolerant systems to prevent failures before they happen.
  • Define SLIs, SLOs, and SLAs, ensuring proactive performance monitoring and incident resolution.
  • Architect and manage scalable cloud infrastructure (AWS) for massive real-time data processing.
  • Optimize containerized environments (Kubernetes, Docker) to support multi-region deployments.
  • Lead the adoption of infrastructure as code (Terraform) to fully automate infrastructure management.
  • Build and refine a world-class monitoring, alerting, and logging system using Prometheus, Grafana, OpenTelemetry, and Datadog.
  • Lead incident response and on-call operations, reducing mean time to detection (MTTD) and mean time to resolution (MTTR).
  • Conduct blameless postmortems and continuously improve system resilience.
  • Reduce manual intervention through automated deployment, scaling, and failover mechanisms.
  • Partner with Security & Compliance teams to ensure infrastructure meets HIPAA and SOC 2 standards.
  • Lead disaster recovery and business continuity planning to ensure critical healthcare services are always available.
  • Drive technical strategy and roadmap for scalability, monitoring, and reliability engineering.
  • Collaborate with Product, Engineering, and Infrastructure teams to align SRE initiatives with business priorities.

What You Have

  • 10+ years of experience in Site Reliability Engineering or Cloud Infrastructure.
  • Proven success scaling high-traffic, mission-critical platforms in SaaS, IoT, or healthcare.
  • Deep expertise in cloud platforms (AWS), Kubernetes, and distributed systems.
  • Strong background in monitoring, logging, and observability with Prometheus, OpenTelemetry, or similar tools.
  • Hands-on experience with incident management, postmortems, and building resilient systems.
  • Deep knowledge of CI/CD automation, GitOps, and infrastructure as code (Terraform, etc.).
  • A mature leadership approach, with the ability to drive technical strategy while growing and mentoring a high-performance SRE team.
  • Strong understanding of network security, access management, and compliance frameworks (HIPAA, SOC 2).

Bonus Points If You Have:

  • Experience with healthcare IT, including EHR data, FHIR, and HL7 interoperability.
  • Expertise in real-time distributed systems, event-driven architectures, or large-scale data pipelines.
  • Prior experience leading on-call rotations and major incident management processes.

Logistics, Perks & Benefits

  • Built for collaboration - our team a hybrid schedule of 3 days/week minimum from our New York City office
  • Equity in a high-growth company scaling toward $400M+ ARR and backed by leading investors
  • Full health, dental, and vision coverage, a 401k, paid time off, paid parental leave and all the tools you need to do your best work
  • Autonomy to solve meaningful problems with work that ships quickly and makes a difference

Compensation

The expected salary range for this role is $200,000 – $250,000 for New York-based candidates. Actual compensation within this range will be determined based on relevant experience, skills, and qualifications. In exceptional cases, where a candidate’s experience or qualifications significantly exceed those anticipated for this role, we may consider the candidate for a more senior level. This role may also be eligible for equity and bonus compensation.
 

HQ

Kontakt.io New York, New York, USA Office

133 W 19th Street , New York, New York , United States, 10011

Similar Jobs

25 Days Ago
Hybrid
New York, NY, USA
Senior level
Senior level
Financial Services
Lead SRE for Sales Execution platforms responsible for stability, availability, resiliency, incident leadership, RCA, and operational maturity. Partner with Front Office, Product, Development, and Infrastructure to drive SRE adoption, observability, automation, and AI-assisted reliability workflows while mentoring engineers and owning outcomes for business-critical services.
Top Skills: AnsibleAWSAzureCi/CdContainersDynatraceEnterprise-Authorized AiGCPGeneosGrafanaItilKubernetesMicroservicesOpenshiftPowershellPythonShellSplunkTerraform
15 Days Ago
Easy Apply
Hybrid
New York, NY, USA
Easy Apply
184K-240K Annually
Senior level
184K-240K Annually
Senior level
Artificial Intelligence • Marketing Tech • Software
Lead SRE to provide technical leadership for a multi-cloud, multi-region active-active content platform. Define automation and IaC strategy, design core platform and logging/observability systems, establish capacity and performance frameworks, lead cross-team reliability initiatives, improve on-call and incident response, mentor engineers, and proactively address systemic reliability and scaling challenges.
Top Skills: Apache KafkaApache PulsarAWSCassandraChefEksGCPGkeGoGrafana AlloyGrafana LokiKubernetesLinuxNode.jsPrometheusPythonRubyScylladbShell ScriptingTempoTerraformThanos
4 Days Ago
In-Office
100K-150K Annually
Senior level
100K-150K Annually
Senior level
Artificial Intelligence • Information Technology • Software • Consulting
Build and maintain enterprise digital experiences on AEM (Sites, Assets, Cloud). Hands-on development of components, templates, workflows, integrations, front-end experiences, GraphQL/headless delivery, CI/CD, performance tuning, and stakeholder communication.
Top Skills: Adobe Experience Manager (Aem)Adobe Experience PlatformAem As A Cloud ServiceAem AssetsAem SitesCi/CdContent FragmentsDispatcherEdge Delivery ServicesFranklinGraphQLHeadless DeliveryHlxHtlJavaJavaScriptMagentoModern Js Build ToolsOsgiSalesforce Commerce CloudSlingSpa Editor

What you need to know about the NYC Tech Scene

As the undisputed financial capital of the world, New York City is an epicenter of startup funding activity. The city has a thriving fintech scene and is a major player in verticals ranging from AI to biotech, cybersecurity and digital media. It also has universities like NYU, Columbia and Cornell Tech attracting students and researchers from across the globe, providing the ecosystem with a constant influx of world-class talent. And its East Coast location and three international airports make it a perfect spot for European companies establishing a foothold in the United States.

Key Facts About NYC Tech

  • Number of Tech Workers: 549,200; 6% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Capgemini, Bloomberg, IBM, Spotify
  • Key Industries: Artificial intelligence, Fintech
  • Funding Landscape: $25.5 billion in venture capital funding in 2024 (Pitchbook)
  • Notable Investors: Greycroft, Thrive Capital, Union Square Ventures, FirstMark Capital, Tiger Global Management, Tribeca Venture Partners, Insight Partners, Two Sigma Ventures
  • Research Centers and Universities: Columbia University, New York University, Fordham University, CUNY, AI Now Institute, Flatiron Institute, C.N. Yang Institute for Theoretical Physics, NASA Space Radiation Laboratory

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account