Maximum of 25 job preferences reached.
Top Reliability Engineer Jobs in NYC, NY
Digital Media • Gaming • Information Technology • Software • Sports • Esports • Big Data Analytics
Lead reliability, scalability, and operational excellence of large-scale database platforms across cloud and on-prem. Build automation-first database infrastructure (Kubernetes operators, IaC, GitOps), drive monitoring/SLOs, incident leadership, performance and cost optimization, and partner with application teams on safe schema/migration practices. Mentor engineers and evaluate AI-assisted workflows to improve productivity and reliability.
Top Skills:
AerospikeArgocdAuroraClaudeCloud SqlCursorDatabase OperatorsEksFluxcdGithub CopilotGitopsGkeGoKubernetesMcpMongoDBMySQLPersistent VolumesPostgresPulumiPythonRedisScylladbStatefulsetsTerraform
Fintech • Information Technology • Software • Financial Services
Own the durability, recoverability, performance, and security of a production PostgreSQL/RDS fleet supporting a live trading platform. Lead replication, failover, backup and restore, disaster-recovery drills, data lifecycle management, access control, encryption, and database observability. Investigate engine-level performance issues including WAL contention, replica lag, bloat, and locking. Build infrastructure and AI-assisted operational tooling while documenting runbooks and reliability decisions.
Top Skills:
AlloydbAmazon AuroraAmazon RdsAWSBashBigQueryElkGrafanaKafkaKubernetesLinuxPostgresPrometheusPythonSQLTerraform
Artificial Intelligence • Software
As a Forward Deployed Reliability Engineer, you ensure stability of workflows, resolve issues swiftly, automate tasks, and drive product improvements through collaboration and documentation.
Top Skills:
JavaPythonSparkSQL
Artificial Intelligence • Software
As a Product Reliability Engineer, you'll ensure service health and performance, tackle outages, and enhance code stability while improving observability and resilience in complex systems.
Top Skills:
CSSDjangoFlaskGoHTMLJavaJavaScriptPrometheusPythonRubyRuby On Rails
Fintech • Software
The Senior Site Reliability Engineer ensures SaaS platforms remain reliable, performant, secure, and scalable. Responsibilities include building cloud infrastructure, implementing monitoring and alerting, automating operational runbooks and deployments, managing Infrastructure as Code, applying AI-powered observability and remediation, supporting Kubernetes and cloud networking, and leading incident triage and root-cause analysis during 24/7 on-call rotations.
Top Skills:
AIAiopsAksAnsibleAppdynamicsAWSAzureAzure DevopsBashC# .NetCi/CdCloud NetworkingCloudopsCosmos DbDatadogDynatraceEksFirewallsHarnessIdera Sql Diagnostic ManagerInfrastructure As CodeJavaJenkinsKubernetesLinuxLoad BalancingNew RelicPowershellPythonRedgate Sql MonitorSolarwinds Database Performance AnalyzerSQLTerraformWindows
Healthtech • Information Technology • Software • Telehealth
Develop, monitor, and maintain distributed production systems and AWS-based microservices infrastructure. Build automation, tooling, and repeatable processes that improve uptime, scalability, security, and operational efficiency. Support product engineering teams with performance, scaling, incident diagnosis, and production debugging. Analyze and tune systems, code, and networking while participating in on-call operations and blameless post-mortems.
Top Skills:
AWSDnsDockerGCPGenaiHttp/HttpsKubernetesLoad BalancersNtpReverse ProxiesTcp/IpTlsWeb Application Firewalls
Software • Defense
Own reliability, scalability, security, observability, and incident response for production applications across AWS and on-premises DoD environments. Build monitoring and alerting, define SLIs and SLOs, lead post-incident reviews, automate infrastructure with Terraform and Ansible, operate Kubernetes clusters, embed RMF and STIG controls, reduce operational toil, and support secure air-gapped deployments.
Top Skills:
AlloyAnsibleAWSAws GovcloudBashDatadogElk StackGithub ActionsGitlab Ci/CdGitopsGoGrafanaHyper-VIstioJenkinsKubernetesLinkerdLokiNutanixPrometheusProxmoxPythonRmfSecurity+StigsTerraformVMware
Blockchain • Fintech • Payments • Financial Services • Cryptocurrency • Web3 • Infrastructure as a Service (IaaS)
Manage AWS and GCP cloud environments, scale infrastructure globally, shape technical architecture, and build reliable CI/CD pipelines. Automate security and compliance controls, improve infrastructure performance, and develop systems interacting with smart contracts across multiple blockchains. The role requires Terraform, shell scripting, GitHub Actions, Docker, production cloud operations, and observability experience, with Kubernetes, networking, and fintech compliance knowledge preferred.
Top Skills:
AWSCi/CdDockerFirewallsGCPGithub ActionsGkeGoHelmInfrastructure As CodeKubernetesLoad BalancersMtlsNode.jsPciShellSsl/TlsTerraformTypescriptVpcZero Trust
9 Days AgoSaved
Easy Apply
Easy Apply
Cloud • Security • Software • Cybersecurity • Automation
Build and operate reliable, scalable production infrastructure for GitLab’s user-facing services. Responsibilities include developing infrastructure automation and tooling, managing Kubernetes deployments, maintaining infrastructure as code, supporting CI/CD and GitOps, participating in on-call and incident response, improving observability and SLOs, troubleshooting production systems, and documenting operational practices. The role spans Intermediate through Senior Staff levels and requires strong software engineering, cloud, reliability, and asynchronous collaboration skills.
Top Skills:
AlertingAWSCi/CdGCPGitopsGoInfrastructure As CodeKubernetesLoggingMetricsRubySlisSlosTerraform
Fintech
Own the enterprise observability strategy and governance, defining SLIs, SLOs, error budgets, KPIs, reliability standards, and monitoring practices. Architect telemetry, tracing, logging, metrics, synthetics, APM, and RUM solutions across complex platforms. Advise engineering, SRE, infrastructure, product, and operations leaders; develop executive service-health reporting; analyze incidents and alerts to improve reliability, detection, and resiliency. Provide technical leadership, mentorship, and standards governance for enterprise observability engineering.
Top Skills:
ApmCloud PlatformsDatadogDistributed TracingDynatraceElasticGrafanaKubernetesMicroservicesNew RelicOpentelemetryPrometheusRumSplunkTelemetry Frameworks
Healthtech • Social Impact • Software
Define and scale reliability practices across the company by creating SLO/SLA frameworks, improving observability, evolving incident response, building self-service tooling and scorecards, and driving cross-team adoption to enable teams to build and operate reliable production systems at scale.
Top Skills:
AWSDatadogEksKubernetesPostgresTerraform
Fintech • Financial Services
Leads Forge’s Site Reliability Engineering team and drives system availability, incident management, observability, disaster recovery, automation, and production support. Partners with Platform, Engineering, Security, Compliance, Risk, and Product teams to improve reliability and operational maturity in a regulated environment. The role combines hands-on technical leadership with hiring, coaching, performance management, architecture contributions, troubleshooting, and development of secure, scalable infrastructure.
Top Skills:
Amazon CloudwatchAnsibleAWSAzureCi/CdDatadogInfrastructure As CodeKubernetesTerraform
New
Track Smarter, Apply Better.
Ditch the spreadsheets. Organize your job search with our freeApplication Tracker.
Use For Free
20 Days AgoSaved
Artificial Intelligence • Cloud • Fintech • Information Technology • Insurance • Financial Services • Big Data Analytics
Leads cloud platform engineering and site reliability initiatives across AWS, GCP, and Azure. Designs reusable Terraform infrastructure, governed platform patterns, CI/CD integrations, and full-stack solutions for applications, data platforms, and AI services. Establishes SRE practices including SLOs, SLIs, error budgets, observability, incident response, and cost optimization. Partners with application, security, architecture, and business teams; mentors engineers, facilitates technical reviews, resolves operational issues, and drives standardization and process improvement.
Top Skills:
.NetAgentic AiAmazon EksAPIsAWSAws CdkAzureC#Ci/CdCloudFormationFinopsGCPGitGithub ActionsHarnessJavaJavaScriptJfrog ArtifactoryKubernetesPowershellPythonRagSlo/SliSQLTerraformYaml
21 Days AgoSaved
Artificial Intelligence • Cloud • Fintech • Information Technology • Insurance • Financial Services • Big Data Analytics
Leads a cloud SRE shared-services organization supporting AWS application platforms. Responsibilities include managing and mentoring engineers, modernizing infrastructure, governing Terraform modules and AWS patterns, building self-service provisioning and Harness CI/CD automation, establishing SLOs and error budgets, resolving critical incidents, implementing observability, improving resiliency, and planning disaster recovery. The leader also drives GenAI adoption, stakeholder alignment, security, governance, cost optimization, and engineering best practices across technology teams.
Top Skills:
Agentic AiAmazon CloudwatchAWSCi/CdGenerative AiHarnessInfrastructure As CodeJfrogKubernetesNew RelicTerraform Hcl
Artificial Intelligence • Legal Tech • Software
The Staff Site Reliability Engineer will architect reliability strategies, manage observability, and drive operational excellence across teams, particularly in large-scale production systems.
Top Skills:
Cloud InfrastructureDistributed Systems
Artificial Intelligence • Legal Tech • Software
As a Senior Site Reliability Engineer, you'll operate foundational platform services, enhance reliability standards, automate processes, and work with engineering teams to improve systems.
Top Skills:
Cloud InfrastructureKubernetesObservability Tools
Aerospace • Artificial Intelligence • Machine Learning • Robotics • Software
Lead operational reliability and platform enablement for Databricks: build monitoring, CI/CD, deployment standards, compute and job policies, observability, runbooks, and governance to support secure, cost-aware, production data workloads across regulated environments. Mentor engineers and align platform with cloud/infrastructure and compliance requirements.
Top Skills:
Ci/CdDatabricksDatabricks Asset BundlesDatabricks WorkflowsDelta LakeInfrastructure-As-CodeService PrincipalsUnity CatalogVersion Control (Git)
Financial Services
Lead SRE for Sales Execution platforms responsible for stability, availability, resiliency, incident leadership, RCA, and operational maturity. Partner with Front Office, Product, Development, and Infrastructure to drive SRE adoption, observability, automation, and AI-assisted reliability workflows while mentoring engineers and owning outcomes for business-critical services.
Top Skills:
AnsibleAWSAzureCi/CdContainersDynatraceEnterprise-Authorized AiGCPGeneosGrafanaItilKubernetesMicroservicesOpenshiftPowershellPythonShellSplunkTerraform
Big Data • Healthtech • HR Tech • Machine Learning • Software • Telehealth • Big Data Analytics
Own the reliability, performance, resilience, observability, and security of AWS and Kubernetes infrastructure supporting products and AI/ML workloads. Define SLOs, lead incident response and root-cause analysis, build Terraform automation, optimize cloud costs, reduce operational toil, and establish deployment standards that help engineers ship reliably. Participate in on-call rotations and maintain HIPAA-compliant infrastructure.
Top Skills:
AWSClaudeDatadogGitlabGoHipaaIstioKubernetesNatsPostgresPythonSoc 2TerraformTypescript
Reposted 26 Days AgoSaved
Cloud • Computer Vision • Information Technology • Sales • Security • Cybersecurity
Lead architecture and implementation of reliability improvements across CrowdStrike's cloud-native platform. Build shared libraries and services, drive observability and SLO practices, perform performance and cost optimization, run resilience engineering and chaos experiments, automate infrastructure-as-code, mentor engineers, and embed with product teams to deliver scalable, highly reliable distributed systems at organizational scale.
Top Skills:
AIAlertingAWSCassandraElasticsearchGCPGoInfrastructure-As-CodeJavaKafkaKotlinKubernetesNode.jsObservability (TracingOciOpensearchProfilingProtobufPythonScalaSlos)
Reposted 26 Days AgoSaved
Easy Apply
Easy Apply
Big Data • Cloud • Software • Database
As a Senior Site Reliability Engineer, you'll design and build complex systems, support Atlas platform operations, automate processes, and ensure high availability of services.
Top Skills:
AWSAzureDnsGCPGoHTTPLinuxPythonRubyTls
Fintech
Leads observability engineering for critical customer journeys and production applications. Defines and governs SLIs, SLOs, error budgets, dashboards, alerts, synthetic monitoring, telemetry standards, and service health reporting. Partners with product, engineering, SRE, and operations teams to ensure production readiness, analyze incidents and telemetry, reduce alert fatigue, improve detection accuracy, and drive reliability improvements. Provides technical leadership and mentorship on monitoring, logging, tracing, metrics, and observability best practices.
Top Skills:
Application Performance MonitoringCloud PlatformsDatadogDistributed SystemsDistributed TracingDynatraceElasticGrafanaKubernetesLoggingMicroservicesNew RelicOpentelemetryPrometheusReal User MonitoringSplunkSynthetic MonitoringTelemetry Frameworks
Information Technology
Owns reliability analysis for mission-critical electrical systems in data centers. Develops electrical BIA and FMEA packages, defines monitoring and maintenance strategies, validates one-line diagrams and redundancy assumptions, and recommends testing, spare-parts, redesign, and risk-reduction actions. Supports root-cause analyses, commissioning, vendor reviews, workshops, and design changes while communicating technical risks and tracking reliability improvements. Requires travel of 25% to 40% and occasional maintenance-window support.
Top Skills:
BatteriesCondition MonitoringElectrical DistributionFmeaGeneratorsNeta TestingOne-Line DiagramsPower SystemsProtection And ControlsSwitchgearTransfer SystemsUps Systems
Reposted 28 Days AgoSaved
Easy Apply
Easy Apply
Big Data • Cloud • Software • Database
Develop and maintain Kubernetes runtime environments, support developers, resolve critical issues, and participate in on-call rotations for production systems.
Top Skills:
AWSAzureCert-ManagerCorednsCrdsCriCsiGatekeeperGCPGoHelmKubernetesKustomizeOperatorsPythonTerraform
20 Hours AgoSaved
Fintech • Financial Services
Lead Site Reliability Engineering initiatives for critical, large-scale financial platforms. Responsibilities include defining SLOs, SLIs, and error budgets; designing resilient distributed systems; developing automation and self-service tooling; improving production readiness through testing and capacity planning; leading complex incident response and blameless post-mortems; and promoting sustainable on-call practices across engineering teams.
Top Skills:
AnsibleAWSAzureCloudFormationCloudwatchDatadogDockerElkGCPGrafanaJavaKubernetesLinuxNode.jsOpentelemetryPrometheusPythonSplunkTerraform
Let Your Resume Do The Work
Upload your resume to be matched with jobs you're a great fit for.
Success! We'll use this to further personalize your experience.
Top NYC, NY Companies Hiring Reliability Engineers
See AllPopular NYC, NY Engineering Job Searches
Engineering Jobs in NYC, NY
.NET Developer Jobs in NYC, NY
Android Developer Jobs in NYC, NY
Associate Software Engineer Jobs in NYC, NY
Automation Engineer Jobs in NYC, NY
AWS Engineer Jobs in NYC, NY
Backend Engineer Jobs in NYC, NY
C# Jobs in NYC, NY
C++ Jobs in NYC, NY
Cloud Engineer Jobs in NYC, NY
Controls Engineer Jobs in NYC, NY
CTO Jobs in NYC, NY
Design Engineer Jobs in NYC, NY
DevOps Engineer Jobs in NYC, NY
DevOps Jobs in NYC, NY
Director of Engineering Jobs in NYC, NY
Electrical Engineering Jobs in NYC, NY
Embedded Software Engineer Jobs in NYC, NY
Engineering Manager Jobs in NYC, NY
Field Engineer Jobs in NYC, NY
Front End Developer Jobs in NYC, NY
Full-Stack Engineer Jobs in NYC, NY
Golang Jobs in NYC, NY
Hardware Engineer Jobs in NYC, NY
Infrastructure Engineer Jobs in NYC, NY
iOS Developer Jobs in NYC, NY
Java Developer Jobs in NYC, NY
Java Full-Stack Engineer Jobs in NYC, NY
Javascript Jobs in NYC, NY
Lead Software Engineer Jobs in NYC, NY
Linux Jobs in NYC, NY
Manufacturing Engineer Jobs in NYC, NY
Mechanical Design Engineer Jobs in NYC, NY
Mechanical Engineering Jobs in NYC, NY
Mechatronics Engineering Jobs in NYC, NY
Network Engineer Jobs in NYC, NY
PHP Developer Jobs in NYC, NY
Platform Engineer Jobs in NYC, NY
Principal Architect Jobs in NYC, NY
Principal Engineer Jobs in NYC, NY
Principal Software Engineer Jobs in NYC, NY
Process Engineer Jobs in NYC, NY
Project Engineer Jobs in NYC, NY
Python Jobs in NYC, NY
QA Engineer Jobs in NYC, NY
QA Jobs in NYC, NY
Quality Assurance Automation Engineer Jobs in NYC, NY
Reliability Engineer Jobs in NYC, NY
Robotics Engineer Jobs in NYC, NY
Ruby Jobs in NYC, NY
Salesforce Developer Jobs in NYC, NY
Scala Jobs in NYC, NY
Security Engineer Jobs in NYC, NY
Software Engineer Jobs in NYC, NY
Solutions Architect Jobs in NYC, NY
Solutions Engineer Jobs in NYC, NY
SRE Engineer Jobs in NYC, NY
Staff Engineer Jobs in NYC, NY
Staff Software Engineer Jobs in NYC, NY
Systems Engineer Jobs in NYC, NY
Vice President Of Engineering Jobs in NYC, NY
Web Developer Jobs in NYC, NY
All Filters
Total selected ()
No Results
No Results















.png)














