Top Reliability Engineer Jobs in NYC, NY

Reposted 13 Days AgoSaved
Hybrid
New York, NY
205K-225K Annually
Senior level
205K-225K Annually
Senior level
Artificial Intelligence • Fintech • Payments • Social Impact • Analytics • Financial Services • Automation
As a Senior SRE, you'll ensure reliable and scalable systems, develop observability solutions and infrastructure as code, and lead incident response efforts.
Top Skills: AWSCloudFormationDatadogElkPrometheusTerraform
8 Days AgoSaved
In-Office
New York, NY
169K-208K Annually
Mid level
169K-208K Annually
Mid level
Artificial Intelligence • Software
Lead reliability engineering for a reference design: build and validate availability and RAM models, run cross-discipline FMEAs, quantify failure rates and redundancy trade-offs, and close the loop by feeding fleet field-failure data back into designs to improve maintainability and uptime.
Top Skills: Availability ModelingData Center TopologiesFmeaRam Modeling SoftwareWeibull Analysis
9 Days AgoSaved
In-Office
New York, NY
173K-224K Annually
Mid level
173K-224K Annually
Mid level
Artificial Intelligence • Software
Own reliability for named customer workloads; debug distributed systems across hardware, fabric, and scheduler; run customer-facing incident communications; convert recurring customer pain into engineering fixes; support large-scale compute customers and push internal teams to resolve root causes.
Top Skills: CloudGpuHpcInfinibandKubernetesNcclRoceSlurm
Reposted 12 Hours AgoSaved
Remote
New York, NY
190K-240K Annually
Senior level
190K-240K Annually
Senior level
Artificial Intelligence • Insurance • Software • Automation
Lead design, automation, and optimization of database infrastructure (PostgreSQL/Aurora). Build monitoring, tuning, and scaling strategies, create automation tooling, drive performance and reliability initiatives, and expand into broader SRE responsibilities to improve availability and system health for a growing SaaS platform.
Top Skills: Amazon AuroraCi/CdDockerJavaScriptKubernetesNode.jsPostgresPrismaRedshiftTerraformTerragruntTypescript
Reposted 12 Hours AgoSaved
In-Office or Remote
New York, NY
119K-178K Annually
Senior level
119K-178K Annually
Senior level
Automotive • Information Technology • Other • Transportation • Energy
Perform RAM and FMECA/FMEA analyses, develop fault trees and reliability predictions, support maintainability and logistics analyses, produce reliability growth test plans, contribute to systems engineering documentation, advise design engineers on R&M shortfalls, and present results to management and clients.
Top Skills: Fault Tree AnalysisFmeaFmecaIntegrated Logistics Support (Ils/Ilsa)Iso-9000Mil-Hdbk-217FRam ModellingRam SoftwareStatistical Methods
Reposted 7 Days AgoSaved
Remote
New York, NY
180K-220K Annually
Senior level
180K-220K Annually
Senior level
Software • Defense
Own reliability, scalability, and security for on-prem and AWS deployments. Build observability (Prometheus/Loki/Grafana/ELK), define SLOs/SLIs, lead incident response and postmortems, automate infrastructure (Terraform/Ansible), operate Kubernetes clusters, embed security/compliance controls, eliminate operational toil, and mentor teams.
Top Skills: AlloyAnsibleAWSAws GovcloudBashCloudFormationDatadogElkGithub ActionsGitlab Ci/CdGoGrafanaJenkinsKubernetesLokiPrometheusPythonRmfStigsTerraform
Reposted 7 Days AgoSaved
Easy Apply
Remote
New York, NY
Easy Apply
146K-225K Annually
Junior
146K-225K Annually
Junior
Big Data • Fintech • Mobile • Payments • Financial Services
Design and build a centralized reliability platform and developer-facing APIs. Implement AI agents for incident triage, log/trace summarization, and recommended actions. Own projects end-to-end and collaborate with product, infra, data, and SRE teams.
Top Skills: Ai FrameworksAPIsClaudeCopilotCursorLlmsPython
Reposted 8 Days AgoSaved
Remote or Hybrid
New York, NY
175K-200K Annually
Senior level
175K-200K Annually
Senior level
eCommerce • Fintech • Payments • Software
The role involves ensuring software reliability and performance, managing incidents, developing infrastructure automation, and mentoring junior engineers within a platform team.
Top Skills: AWSCloudFormationDatadogKubernetesOpentelemetryRubyRuby On RailsTerraform
9 Days AgoSaved
Easy Apply
Remote
New York, NY
Easy Apply
100K-110K Annually
Mid level
100K-110K Annually
Mid level
Healthtech • Software
Operate and maintain AWS-hosted MERN applications and large-scale data workflows. Manage serverless and Spark-based pipelines, perform incident response and on-call duties, engineer automation to eliminate operational toil, ensure HIPAA/SOC2/HITRUST compliance, build observability and lead blameless post-mortems.
Top Skills: Amazon EcsAmazon EksAmazon EmrAthenaAws GlueAws LambdaAws SnsAws SqsCloudwatchEc2IamJavaScriptMernMySQLNode.jsOpentofuPysparkPythonRabbitMQTerraformTypescriptVpc
Reposted 18 Days AgoSaved
In-Office
New York, NY
160K-300K Annually
Senior level
160K-300K Annually
Senior level
Artificial Intelligence • Machine Learning • Natural Language Processing • Software • Financial Services • Generative AI
As a Site Reliability Engineer, you'll design and improve critical production systems, lead incident response, and enhance observability while embedding with product teams to ensure reliability and performance at scale.
Top Skills: AWSC++Ci/CdGoPythonRust
Reposted 19 Days AgoSaved
Easy Apply
Hybrid
New York, NY
Easy Apply
111K-218K Annually
Mid level
111K-218K Annually
Mid level
Big Data • Cloud • Software • Database
The Site Reliability Engineer designs and builds infrastructure for a global cloud service, implements automation, and optimizes system performance while managing on-call operations.
Top Skills: AWSDnsGCPHTTPKubernetesLinuxAzureProgramming LanguagesTls
Reposted 11 Days AgoSaved
Easy Apply
Remote or Hybrid
New York, NY
Easy Apply
200K-230K Annually
Senior level
200K-230K Annually
Senior level
Artificial Intelligence • Machine Learning
Lead development of AI-assisted reliability tooling, own incident response end-to-end, improve observability and SLO/SLI frameworks, scale single-tenant SaaS operations, mentor engineers, and reduce recurring operational toil through engineering and automation.
Top Skills: Cloud PlatformsGoKubernetesLinuxLlm/Ai ToolingLogs And TracingObservability ToolingPythonSlo/Sli Frameworks
New

Cut your apply time in half.

Use ourAI Assistantto automatically fill your job applications.

Use For Free
Application Tracker Preview
Reposted 8 Days AgoSaved
Remote
New York, NY
145K-180K Annually
Senior level
145K-180K Annually
Senior level
Legal Tech • Software
Lead automation and optimization of Filevine's data platform: performance tune MSSQL/Postgres, optimize Snowflake, provision infrastructure with Terraform/AWS, run stateful containers on Kubernetes, integrate AI/LLM and MCP for operational automation, manage CI/CD, capacity planning, documentation, and serve in 24/7 on-call rotation.
Top Skills: AWSC#DapperDockerDynamoDBEntity FrameworkGitlabKubernetesLlmsMcp (Model Context Protocol)Microsoft Sql Server (Mssql)Octopus DeployOpensearchPostgresPowershellPythonRedisSnowflakeTerraform
Reposted 23 Days AgoSaved
Easy Apply
Remote or Hybrid
New York, NY
Easy Apply
127K-249K Annually
Senior level
127K-249K Annually
Senior level
Big Data • Cloud • Software • Database
As a Senior Site Reliability Engineer, you'll design and build complex systems, support Atlas platform operations, automate processes, and ensure high availability of services.
Top Skills: AWSAzureDnsGCPGoHTTPLinuxPythonRubyTls
Reposted 25 Days AgoSaved
In-Office
New York, NY
100K-250K Annually
Mid level
100K-250K Annually
Mid level
Fintech • Payments • Financial Services
Design, build, and maintain highly reliable, observable production services. Automate operations, performance-tune cloud deployments, own incident response/on-call, mentor engineers, and improve system reliability and scalability.
Top Skills: AWSAzureDatadogDockerEc2GCPGoKubernetesRustTerraform
16 Days AgoSaved
Remote
New York, NY
150K-200K Annually
Senior level
150K-200K Annually
Senior level
Artificial Intelligence • Cloud • Software • Infrastructure as a Service (IaaS)
Ensure stability and resilience of Runpod's distributed AI platform by defining SLIs/SLOs, leading incident response, building observability and reliability tooling, automating operational workflows, and partnering with engineering teams to reduce toil and improve production readiness.
Top Skills: BashCi/CdContainerized Production SystemsGoGpu Observability ToolingGrafanaInfrastructure As CodeLinuxPrometheusPython
Reposted 25 Days AgoSaved
In-Office
New York, NY
273K-369K Annually
Senior level
273K-369K Annually
Senior level
Artificial Intelligence • Legal Tech • Software
The Staff Site Reliability Engineer will architect reliability strategies, manage observability, and drive operational excellence across teams, particularly in large-scale production systems.
Top Skills: Cloud InfrastructureDistributed Systems
Reposted 25 Days AgoSaved
In-Office
New York, NY
237K-321K Annually
Senior level
237K-321K Annually
Senior level
Artificial Intelligence • Legal Tech • Software
As a Senior Site Reliability Engineer, you'll operate foundational platform services, enhance reliability standards, automate processes, and work with engineering teams to improve systems.
Top Skills: Cloud InfrastructureKubernetesObservability Tools
Reposted 16 Days AgoSaved
Easy Apply
Remote or Hybrid
New York, NY
Easy Apply
Internship
Internship
Cloud • Information Technology • Security • Software • Cybersecurity
This internship role focuses on SRE skills, requiring collaboration and problem-solving in dynamic environments for Zscaler's Zero Trust Exchange team.
Top Skills: AnsibleAws EcsKubernetesLinuxPythonTerraform
Reposted 17 Days AgoSaved
Easy Apply
Remote
New York, NY
Easy Apply
Mid level
Mid level
Cloud • Security • Software • Cybersecurity • Automation
As a Cloud Cost Utilization SRE at GitLab, you'll manage cloud spending, improve tracking and optimization of cloud usage, and collaborate with finance and engineering teams to enhance cost efficiency across AWS and GCP.
Top Skills: AnsibleAWSElkGCPGrafanaLokiMimirPrometheusTempoTerraform
Reposted 14 Days AgoSaved
Remote
New York, NY
75K-150K Annually
Senior level
75K-150K Annually
Senior level
Database • Analytics
As a Database Reliability Engineer at ClickHouse, you'll improve reliability, manage escalation processes, support incident response, and enhance database performance while collaborating across teams.
Top Skills: AWSAzureC++ClickhouseGoogle Cloud PlatformPythonShellSQL
Reposted 14 Days AgoSaved
Remote or Hybrid
New York, NY
Senior level
Senior level
Software
Lead reliability engineering for Silicon Photonics hardware: define and validate reliability models, perform MTBF/MTBCF predictions, analyze field data, direct verification testing and root-cause analysis, drive corrective actions, and mentor cross-functional teams to improve product reliability.
Top Skills: Derating AnalysisDfmeaMtbcfMtbfSherlockSilicon PhotonicsTelcordiaThermal DesignWindchill Qs
Reposted 20 Days AgoSaved
Easy Apply
Remote or Hybrid
New York, NY
Easy Apply
227K-272K Annually
Senior level
227K-272K Annually
Senior level
eCommerce • Healthtech • Kids + Family • Retail • Social Media
Own and evolve Babylist's AWS infrastructure and developer platform using Terraform and Kubernetes. Improve CI/CD reliability, support engineers across environments, define monitoring and alerting standards, lead incident response and postmortems, and shape platform architecture to scale for millions of users.
Top Skills: AWSCdnCircleCICronitorDatadogDnsEksGithub ActionsKubernetesLoad BalancersMySQLPagerdutyRdsRedisRuby On RailsSentrySidekiqTerraform
Reposted 15 Days AgoSaved
Remote
New York, NY
Mid level
Mid level
Information Technology • Software • Database • Automation
Owner of on-prem reliability and escalations: reproduce and resolve L2/L3 issues across heterogeneous Kubernetes environments, build diagnostics and automation, improve CI and e2e test stability, establish performance baselines, harden install/upgrade flows, and write tooling in Python/Go/Rust to reduce repeat incidents.
Top Skills: BenchmarkingCiCi/CdContainersE2E TestingGoHealth ChecksHelmInstallersIntegration TestingKubernetesLoad GenerationLogsMetricsNetworkingObservabilityPackagingProfilingPythonRbacRustStorageSupport BundlesTraces
3 Days AgoSaved
In-Office
New York, NY
125K-150K Annually
Junior
125K-150K Annually
Junior
Information Technology • Financial Services
Operate and maintain low-latency electronic trading systems across global venues; troubleshoot incidents, support external trading connections, perform system triage/optimization, collaborate with traders and exchanges, and manage deployments in Windows/Linux environments.
Top Skills: AnsibleBashEsmGoLinuxOmsPerlPuppetPythonSaltSQLTcpUdpWindows
All Filters
JobType
New Jobs
Job Category
Experience
Industry
Company Name
Company Size

Sign up now Access later

Create Free Account