Top Reliability Engineer Jobs in NYC, NY

Reposted 7 Days AgoSaved
Remote or Hybrid
New York, NY
168K-210K Annually
Senior level
168K-210K Annually
Senior level
Digital Media • Gaming • Information Technology • Software • Sports • Esports • Big Data Analytics
Lead reliability, scalability, and operational excellence of large-scale database platforms across cloud and on-prem. Build automation-first database infrastructure (Kubernetes operators, IaC, GitOps), drive monitoring/SLOs, incident leadership, performance and cost optimization, and partner with application teams on safe schema/migration practices. Mentor engineers and evaluate AI-assisted workflows to improve productivity and reliability.
Top Skills: AerospikeArgocdAuroraClaudeCloud SqlCursorDatabase OperatorsEksFluxcdGithub CopilotGitopsGkeGoKubernetesMcpMongoDBMySQLPersistent VolumesPostgresPulumiPythonRedisScylladbStatefulsetsTerraform
8 Days AgoSaved
Easy Apply
Remote or Hybrid
New York, NY
Easy Apply
225K-265K Annually
Senior level
225K-265K Annually
Senior level
Fintech • Information Technology • Software • Financial Services
Own the durability, recoverability, performance, and security of a production PostgreSQL/RDS fleet supporting a live trading platform. Lead replication, failover, backup and restore, disaster-recovery drills, data lifecycle management, access control, encryption, and database observability. Investigate engine-level performance issues including WAL contention, replica lag, bloat, and locking. Build infrastructure and AI-assisted operational tooling while documenting runbooks and reliability decisions.
Top Skills: AlloydbAmazon AuroraAmazon RdsAWSBashBigQueryElkGrafanaKafkaKubernetesLinuxPostgresPrometheusPythonSQLTerraform
Reposted One Month AgoSaved
Hybrid
New York, NY
110K-147K Annually
Entry level
110K-147K Annually
Entry level
Artificial Intelligence • Software
As a Forward Deployed Reliability Engineer, you ensure stability of workflows, resolve issues swiftly, automate tasks, and drive product improvements through collaboration and documentation.
Top Skills: JavaPythonSparkSQL
Reposted One Month AgoSaved
Hybrid
New York, NY
96K-140K Annually
Mid level
96K-140K Annually
Mid level
Artificial Intelligence • Software
As a Product Reliability Engineer, you'll ensure service health and performance, tackle outages, and enhance code stability while improving observability and resilience in complex systems.
Top Skills: CSSDjangoFlaskGoHTMLJavaJavaScriptPrometheusPythonRubyRuby On Rails
YesterdaySaved
Remote or Hybrid
New York, NY
Senior level
Senior level
Fintech • Software
The Senior Site Reliability Engineer ensures SaaS platforms remain reliable, performant, secure, and scalable. Responsibilities include building cloud infrastructure, implementing monitoring and alerting, automating operational runbooks and deployments, managing Infrastructure as Code, applying AI-powered observability and remediation, supporting Kubernetes and cloud networking, and leading incident triage and root-cause analysis during 24/7 on-call rotations.
Top Skills: AIAiopsAksAnsibleAppdynamicsAWSAzureAzure DevopsBashC# .NetCi/CdCloud NetworkingCloudopsCosmos DbDatadogDynatraceEksFirewallsHarnessIdera Sql Diagnostic ManagerInfrastructure As CodeJavaJenkinsKubernetesLinuxLoad BalancingNew RelicPowershellPythonRedgate Sql MonitorSolarwinds Database Performance AnalyzerSQLTerraformWindows
2 Days AgoSaved
Easy Apply
Remote or Hybrid
New York, NY
Easy Apply
180K-220K Annually
Senior level
180K-220K Annually
Senior level
Healthtech • Information Technology • Software • Telehealth
Develop, monitor, and maintain distributed production systems and AWS-based microservices infrastructure. Build automation, tooling, and repeatable processes that improve uptime, scalability, security, and operational efficiency. Support product engineering teams with performance, scaling, incident diagnosis, and production debugging. Analyze and tune systems, code, and networking while participating in on-call operations and blameless post-mortems.
Top Skills: AWSDnsDockerGCPGenaiHttp/HttpsKubernetesLoad BalancersNtpReverse ProxiesTcp/IpTlsWeb Application Firewalls
3 Days AgoSaved
Remote
New York, NY
180K-220K Annually
Senior level
180K-220K Annually
Senior level
Software • Defense
Own reliability, scalability, security, observability, and incident response for production applications across AWS and on-premises DoD environments. Build monitoring and alerting, define SLIs and SLOs, lead post-incident reviews, automate infrastructure with Terraform and Ansible, operate Kubernetes clusters, embed RMF and STIG controls, reduce operational toil, and support secure air-gapped deployments.
Top Skills: AlloyAnsibleAWSAws GovcloudBashDatadogElk StackGithub ActionsGitlab Ci/CdGitopsGoGrafanaHyper-VIstioJenkinsKubernetesLinkerdLokiNutanixPrometheusProxmoxPythonRmfSecurity+StigsTerraformVMware
3 Days AgoSaved
Remote
New York, NY
170K-225K Annually
Mid level
170K-225K Annually
Mid level
Blockchain • Fintech • Payments • Financial Services • Cryptocurrency • Web3 • Infrastructure as a Service (IaaS)
Manage AWS and GCP cloud environments, scale infrastructure globally, shape technical architecture, and build reliable CI/CD pipelines. Automate security and compliance controls, improve infrastructure performance, and develop systems interacting with smart contracts across multiple blockchains. The role requires Terraform, shell scripting, GitHub Actions, Docker, production cloud operations, and observability experience, with Kubernetes, networking, and fintech compliance knowledge preferred.
Top Skills: AWSCi/CdDockerFirewallsGCPGithub ActionsGkeGoHelmInfrastructure As CodeKubernetesLoad BalancersMtlsNode.jsPciShellSsl/TlsTerraformTypescriptVpcZero Trust
9 Days AgoSaved
Easy Apply
Remote
New York, NY
Easy Apply
Senior level
Senior level
Cloud • Security • Software • Cybersecurity • Automation
Build and operate reliable, scalable production infrastructure for GitLab’s user-facing services. Responsibilities include developing infrastructure automation and tooling, managing Kubernetes deployments, maintaining infrastructure as code, supporting CI/CD and GitOps, participating in on-call and incident response, improving observability and SLOs, troubleshooting production systems, and documenting operational practices. The role spans Intermediate through Senior Staff levels and requires strong software engineering, cloud, reliability, and asynchronous collaboration skills.
Top Skills: AlertingAWSCi/CdGCPGitopsGoInfrastructure As CodeKubernetesLoggingMetricsRubySlisSlosTerraform
14 Days AgoSaved
In-Office
New York, NY
98K-116K Annually
Senior level
98K-116K Annually
Senior level
Fintech
Own the enterprise observability strategy and governance, defining SLIs, SLOs, error budgets, KPIs, reliability standards, and monitoring practices. Architect telemetry, tracing, logging, metrics, synthetics, APM, and RUM solutions across complex platforms. Advise engineering, SRE, infrastructure, product, and operations leaders; develop executive service-health reporting; analyze incidents and alerts to improve reliability, detection, and resiliency. Provide technical leadership, mentorship, and standards governance for enterprise observability engineering.
Top Skills: ApmCloud PlatformsDatadogDistributed TracingDynatraceElasticGrafanaKubernetesMicroservicesNew RelicOpentelemetryPrometheusRumSplunkTelemetry Frameworks
Reposted One Month AgoSaved
Hybrid
New York, NY
182K-250K Annually
Senior level
182K-250K Annually
Senior level
Healthtech • Social Impact • Software
Define and scale reliability practices across the company by creating SLO/SLA frameworks, improving observability, evolving incident response, building self-service tooling and scorecards, and driving cross-team adoption to enable teams to build and operate reliable production systems at scale.
Top Skills: AWSDatadogEksKubernetesPostgresTerraform
20 Days AgoSaved
Easy Apply
Hybrid
New York, NY
Easy Apply
150K-220K Annually
Expert/Leader
150K-220K Annually
Expert/Leader
Fintech • Financial Services
Leads Forge’s Site Reliability Engineering team and drives system availability, incident management, observability, disaster recovery, automation, and production support. Partners with Platform, Engineering, Security, Compliance, Risk, and Product teams to improve reliability and operational maturity in a regulated environment. The role combines hands-on technical leadership with hiring, coaching, performance management, architecture contributions, troubleshooting, and development of secure, scalable infrastructure.
Top Skills: Amazon CloudwatchAnsibleAWSAzureCi/CdDatadogInfrastructure As CodeKubernetesTerraform
New

Track Smarter, Apply Better.

Ditch the spreadsheets. Organize your job search with our freeApplication Tracker.

Use For Free
Application Tracker Preview
20 Days AgoSaved
Hybrid
New York, NY
148K-211K Annually
Senior level
148K-211K Annually
Senior level
Artificial Intelligence • Cloud • Fintech • Information Technology • Insurance • Financial Services • Big Data Analytics
Leads cloud platform engineering and site reliability initiatives across AWS, GCP, and Azure. Designs reusable Terraform infrastructure, governed platform patterns, CI/CD integrations, and full-stack solutions for applications, data platforms, and AI services. Establishes SRE practices including SLOs, SLIs, error budgets, observability, incident response, and cost optimization. Partners with application, security, architecture, and business teams; mentors engineers, facilitates technical reviews, resolves operational issues, and drives standardization and process improvement.
Top Skills: .NetAgentic AiAmazon EksAPIsAWSAws CdkAzureC#Ci/CdCloudFormationFinopsGCPGitGithub ActionsHarnessJavaJavaScriptJfrog ArtifactoryKubernetesPowershellPythonRagSlo/SliSQLTerraformYaml
21 Days AgoSaved
Hybrid
New York, NY
185K-265K Annually
Senior level
185K-265K Annually
Senior level
Artificial Intelligence • Cloud • Fintech • Information Technology • Insurance • Financial Services • Big Data Analytics
Leads a cloud SRE shared-services organization supporting AWS application platforms. Responsibilities include managing and mentoring engineers, modernizing infrastructure, governing Terraform modules and AWS patterns, building self-service provisioning and Harness CI/CD automation, establishing SLOs and error budgets, resolving critical incidents, implementing observability, improving resiliency, and planning disaster recovery. The leader also drives GenAI adoption, stakeholder alignment, security, governance, cost optimization, and engineering best practices across technology teams.
Top Skills: Agentic AiAmazon CloudwatchAWSCi/CdGenerative AiHarnessInfrastructure As CodeJfrogKubernetesNew RelicTerraform Hcl
Reposted 22 Days AgoSaved
In-Office
New York, NY
273K-369K Annually
Senior level
273K-369K Annually
Senior level
Artificial Intelligence • Legal Tech • Software
The Staff Site Reliability Engineer will architect reliability strategies, manage observability, and drive operational excellence across teams, particularly in large-scale production systems.
Top Skills: Cloud InfrastructureDistributed Systems
Reposted 22 Days AgoSaved
In-Office
New York, NY
237K-321K Annually
Senior level
237K-321K Annually
Senior level
Artificial Intelligence • Legal Tech • Software
As a Senior Site Reliability Engineer, you'll operate foundational platform services, enhance reliability standards, automate processes, and work with engineering teams to improve systems.
Top Skills: Cloud InfrastructureKubernetesObservability Tools
One Month AgoSaved
Remote
New York, NY
180K-270K Annually
Senior level
180K-270K Annually
Senior level
Aerospace • Artificial Intelligence • Machine Learning • Robotics • Software
Lead operational reliability and platform enablement for Databricks: build monitoring, CI/CD, deployment standards, compute and job policies, observability, runbooks, and governance to support secure, cost-aware, production data workloads across regulated environments. Mentor engineers and align platform with cloud/infrastructure and compliance requirements.
Top Skills: Ci/CdDatabricksDatabricks Asset BundlesDatabricks WorkflowsDelta LakeInfrastructure-As-CodeService PrincipalsUnity CatalogVersion Control (Git)
Reposted 23 Days AgoSaved
Hybrid
New York, NY
Senior level
Senior level
Financial Services
Lead SRE for Sales Execution platforms responsible for stability, availability, resiliency, incident leadership, RCA, and operational maturity. Partner with Front Office, Product, Development, and Infrastructure to drive SRE adoption, observability, automation, and AI-assisted reliability workflows while mentoring engineers and owning outcomes for business-critical services.
Top Skills: AnsibleAWSAzureCi/CdContainersDynatraceEnterprise-Authorized AiGCPGeneosGrafanaItilKubernetesMicroservicesOpenshiftPowershellPythonShellSplunkTerraform
15 Days AgoSaved
Easy Apply
Remote
New York, NY
Easy Apply
191K-226K Annually
Senior level
191K-226K Annually
Senior level
Big Data • Healthtech • HR Tech • Machine Learning • Software • Telehealth • Big Data Analytics
Own the reliability, performance, resilience, observability, and security of AWS and Kubernetes infrastructure supporting products and AI/ML workloads. Define SLOs, lead incident response and root-cause analysis, build Terraform automation, optimize cloud costs, reduce operational toil, and establish deployment standards that help engineers ship reliably. Participate in on-call rotations and maintain HIPAA-compliant infrastructure.
Top Skills: AWSClaudeDatadogGitlabGoHipaaIstioKubernetesNatsPostgresPythonSoc 2TerraformTypescript
Reposted 26 Days AgoSaved
Hybrid
New York, NY
140K-215K Annually
Senior level
140K-215K Annually
Senior level
Cloud • Computer Vision • Information Technology • Sales • Security • Cybersecurity
Lead architecture and implementation of reliability improvements across CrowdStrike's cloud-native platform. Build shared libraries and services, drive observability and SLO practices, perform performance and cost optimization, run resilience engineering and chaos experiments, automate infrastructure-as-code, mentor engineers, and embed with product teams to deliver scalable, highly reliable distributed systems at organizational scale.
Top Skills: AIAlertingAWSCassandraElasticsearchGCPGoInfrastructure-As-CodeJavaKafkaKotlinKubernetesNode.jsObservability (TracingOciOpensearchProfilingProtobufPythonScalaSlos)
Reposted 26 Days AgoSaved
Easy Apply
Remote or Hybrid
New York, NY
Easy Apply
127K-249K Annually
Senior level
127K-249K Annually
Senior level
Big Data • Cloud • Software • Database
As a Senior Site Reliability Engineer, you'll design and build complex systems, support Atlas platform operations, automate processes, and ensure high availability of services.
Top Skills: AWSAzureDnsGCPGoHTTPLinuxPythonRubyTls
21 Days AgoSaved
In-Office
New York, NY
98K-116K Annually
Senior level
98K-116K Annually
Senior level
Fintech
Leads observability engineering for critical customer journeys and production applications. Defines and governs SLIs, SLOs, error budgets, dashboards, alerts, synthetic monitoring, telemetry standards, and service health reporting. Partners with product, engineering, SRE, and operations teams to ensure production readiness, analyze incidents and telemetry, reduce alert fatigue, improve detection accuracy, and drive reliability improvements. Provides technical leadership and mentorship on monitoring, logging, tracing, metrics, and observability best practices.
Top Skills: Application Performance MonitoringCloud PlatformsDatadogDistributed SystemsDistributed TracingDynatraceElasticGrafanaKubernetesLoggingMicroservicesNew RelicOpentelemetryPrometheusReal User MonitoringSplunkSynthetic MonitoringTelemetry Frameworks
12 Days AgoSaved
Remote
New York, NY
Mid level
Mid level
Information Technology
Owns reliability analysis for mission-critical electrical systems in data centers. Develops electrical BIA and FMEA packages, defines monitoring and maintenance strategies, validates one-line diagrams and redundancy assumptions, and recommends testing, spare-parts, redesign, and risk-reduction actions. Supports root-cause analyses, commissioning, vendor reviews, workshops, and design changes while communicating technical risks and tracking reliability improvements. Requires travel of 25% to 40% and occasional maintenance-window support.
Top Skills: BatteriesCondition MonitoringElectrical DistributionFmeaGeneratorsNeta TestingOne-Line DiagramsPower SystemsProtection And ControlsSwitchgearTransfer SystemsUps Systems
Reposted 28 Days AgoSaved
Easy Apply
Remote or Hybrid
New York, NY
Easy Apply
127K-249K Annually
Senior level
127K-249K Annually
Senior level
Big Data • Cloud • Software • Database
Develop and maintain Kubernetes runtime environments, support developers, resolve critical issues, and participate in on-call rotations for production systems.
Top Skills: AWSAzureCert-ManagerCorednsCrdsCriCsiGatekeeperGCPGoHelmKubernetesKustomizeOperatorsPythonTerraform
Expert/Leader
Fintech • Financial Services
Lead Site Reliability Engineering initiatives for critical, large-scale financial platforms. Responsibilities include defining SLOs, SLIs, and error budgets; designing resilient distributed systems; developing automation and self-service tooling; improving production readiness through testing and capacity planning; leading complex incident response and blameless post-mortems; and promoting sustainable on-call practices across engineering teams.
Top Skills: AnsibleAWSAzureCloudFormationCloudwatchDatadogDockerElkGCPGrafanaJavaKubernetesLinuxNode.jsOpentelemetryPrometheusPythonSplunkTerraform
All Filters
JobType
New Jobs
Job Category
Experience
Industry
Company Name
Company Size

Sign up now Access later

Create Free Account