Top Reliability Engineer Jobs in NYC, NY

10 Days AgoSaved
In-Office
New York, NY
350K-475K Annually
Mid level
350K-475K Annually
Mid level
Artificial Intelligence • Information Technology
Own the reliability, performance, and uptime of large-scale post-training and reinforcement learning systems. Debug distributed failures across accelerators, networking, storage, schedulers, and training frameworks; build monitoring, alerting, recovery, checkpointing, and scheduling tools; improve cluster utilization and fault tolerance; support production model runs through on-call rotations and postmortems; and partner closely with research teams during active training runs.
Top Skills: C++DpoGoGpuInfinibandKubernetesLinuxNcclPpoPythonPyTorchRayRdmaRlhfSlurmTpu
10 Days AgoSaved
In-Office
New York, NY
350K-475K Annually
Entry level
350K-475K Annually
Entry level
Artificial Intelligence • Information Technology
Owns end-to-end reliability for Tinker, including CI/CD, observability, service-level objectives, incident response, multi-tenant isolation, resource scheduling, and vulnerability remediation. The role designs monitoring for distributed training systems, improves recovery and checkpointing, and operates Kubernetes clusters supporting heterogeneous GPU workloads. It collaborates closely with platform, security, engineering, and research teams to improve production resilience and utilization.
Top Skills: Ci/CdCloud InfrastructureDistributed SystemsGpu WorkloadsKubernetesLoraProduction Observability
Reposted 14 Hours AgoSaved
Remote
New York, NY
170K-190K Annually
Senior level
170K-190K Annually
Senior level
Artificial Intelligence • HR Tech • Professional Services
Maintain and support platform environments across the software development lifecycle. Configure monitoring, alerting, and self-healing solutions; troubleshoot application and infrastructure issues with developers and testers; maintain technical documentation; design scalable architectural solutions; and mentor junior engineers. The role requires cloud, Kubernetes, operating system, database, scripting, configuration management, .NET, and AI/ML automation expertise.
Top Skills: .NetAi/MlAnsibleAWSAzureBashGCPHelmJenkinsKubernetesLinuxOracle CloudOracle DatabasePowershellService MeshSQL ServerWindows
YesterdaySaved
In-Office or Remote
New York, NY
146K-264K Annually
Senior level
146K-264K Annually
Senior level
Cloud • Security • Software • Cybersecurity
Architect, develop, test, and distribute software, services, and infrastructure supporting Akamai’s cloud hypervisor platforms. Improve observability, automate infrastructure processes, troubleshoot complex distributed-system issues, mentor engineers, and participate in on-call service restoration. The role requires deep Linux, kernel, virtualization, ARM hardware, large-scale infrastructure, DevOps, and configuration-management expertise.
Top Skills: AnsibleArmDevOpsDistributed SystemsKvm/QemuLinuxLinux KernelNested VirtualizationNvidia GraceObservability InfrastructureSaltstack
YesterdaySaved
In-Office or Remote
New York, NY
169K-305K Annually
Expert/Leader
169K-305K Annually
Expert/Leader
Cloud • Security • Software • Cybersecurity
Architect, build, and support reliable network infrastructure and automation for Akamai’s distributed cloud platform. Develop Bash and Python tooling, establish deployment standards, define SLOs, mentor engineers, and troubleshoot complex network issues. The role requires expertise in large-scale distributed systems, TCP/IP, BGP, Linux networking, configuration management, CI/CD, and open-source networking software, with participation in on-call rotations.
Top Skills: AnsibleArgocdBashBgpBirdChefFirewallsFrrGithub ActionsGoGobgpJenkinsLinux NetworkingLoad BalancingPuppetPythonRustSaltstackSlack BotsTcp/Ip
YesterdaySaved
Remote
New York, NY
Senior level
Senior level
Software
Own reliability, performance, scalability, and operational standards across on-premises, private-cloud, and AWS environments. Build infrastructure-as-code, deployment automation, monitoring, alerting, and observability tooling; define SLOs, SLIs, error budgets, and readiness standards. Lead incident response, root-cause analysis, disaster-recovery readiness, release coordination, and preventive automation. Mentor engineers, coach teams on operational practices, coordinate on-call coverage, and ensure infrastructure meets security and compliance requirements.
Top Skills: AnsibleAuto ScalingAWSBashCi/CdClaudeCloudwatchCrowdstrikeDatadogDistributed SystemsDnsDockerEc2EcsGitGithub CopilotGitlabGitopsIamKubernetesLinuxLoad BalancingOracle LinuxPythonQualysRapid7RhelS3Tcp/IpTerraformVpcWireshark
Reposted YesterdaySaved
Remote
New York, NY
160K-208K Annually
Senior level
160K-208K Annually
Senior level
Healthtech • Software
Design, automate, and maintain scalable infrastructure and SRE tooling. Manage Kubernetes clusters, CI/CD, monitoring, and incident response. Improve processes, reduce toil via automation, and collaborate with engineering and data teams to support domestic and international workloads.
Top Skills: AWSAzureContainerdDnsDockerFirewallsGCPGoGrpcHelmKubernetesLinuxLoad BalancingPrometheusPythonRoutingShell ScriptingTcp/IpUdp
2 Days AgoSaved
Remote
New York, NY
Senior level
Senior level
Cloud • Software
The Senior Site Reliability Engineer improves production reliability, resilience, and system availability. Responsibilities include operating Linux and Kubernetes infrastructure, automating workflows, developing tools, managing observability, responding to incidents, participating in on-call rotations, writing runbooks, and conducting blameless postmortems. The role collaborates with distributed teams and stakeholders, supports cloud infrastructure, troubleshoots live systems, and promotes DevOps and SRE best practices.
Top Skills: AWSBare-Metal InfrastructureBashCassandraCi/CdClickhouseGitGoKubernetesLinuxLogsMetricsObservabilityPostgresPythonTerraformTraces
Reposted One Month AgoSaved
Easy Apply
Hybrid
New York, NY
Easy Apply
184K-240K Annually
Senior level
184K-240K Annually
Senior level
Artificial Intelligence • Marketing Tech • Software
Lead SRE to provide technical leadership for a multi-cloud, multi-region active-active content platform. Define automation and IaC strategy, design core platform and logging/observability systems, establish capacity and performance frameworks, lead cross-team reliability initiatives, improve on-call and incident response, mentor engineers, and proactively address systemic reliability and scaling challenges.
Top Skills: Apache KafkaApache PulsarAWSCassandraChefEksGCPGkeGoGrafana AlloyGrafana LokiKubernetesLinuxNode.jsPrometheusPythonRubyScylladbShell ScriptingTempoTerraformThanos
Reposted 2 Days AgoSaved
Remote
New York, NY
Senior level
Senior level
Software
Owns reliability, observability, performance, and security for a multi-region SaaS platform. Responsibilities include managing Datadog, implementing APM and tracing, defining SLOs, developing automation, expanding infrastructure as code and CI/CD, automating operational workflows, maintaining security controls, participating in incident response, documenting procedures, and mentoring engineers.
Top Skills: ApmAzure DevopsAzure Kubernetes ServiceAzure SqlBashBicepCi/CdCosmos DbDatadogDistributed TracingHelmInfrastructure As CodeKey VaultKubernetesKustomizeManaged IdentitiesAzureMicrosoft Entra IdPowershellPythonRedisService BusTerraform
2 Days AgoSaved
Remote
New York, NY
Entry level
Entry level
Artificial Intelligence • Hardware • Software • Semiconductor
Develop kernel-centric reliability solutions for AI compute clusters and production services. Responsibilities include debugging failures, building diagnostic tools, supporting incident response, performing root-cause analysis, improving kernel and software reliability, and collaborating with systems, hardware, ASIC, and architecture teams on reliability-focused designs.
Top Skills: CC++Core Dump HandlingDebuggersDistributed ProgrammingGpusParallel ProgrammingProfilersPythonSanitizersTracing
Reposted 2 Days AgoSaved
Remote
New York, NY
124K-171K Annually
Senior level
124K-171K Annually
Senior level
Healthtech • Pharmaceutical • Manufacturing
Support operational deployments and maintenance for Core Speech production systems. Implement and maintain monitoring, alerting, performance reporting, and capacity planning. Participate in on-call rotation for off-hours reliability support. Drive improvements to system performance and architecture, and collaborate on projects while complying with corporate and quality standards.
New

Track Smarter, Apply Better.

Ditch the spreadsheets. Organize your job search with our freeApplication Tracker.

Use For Free
Application Tracker Preview
Reposted One Month AgoSaved
Easy Apply
Remote or Hybrid
New York, NY
Easy Apply
127K-249K Annually
Senior level
127K-249K Annually
Senior level
Big Data • Cloud • Software • Database
Maintain and improve multi-cloud Kubernetes infrastructure, CI/CD (Argo Workflows/ArgoCD), observability, and networking. Build reliable continuous deployment tooling and onboarding flows, provide internal support, collaborate across Platform Engineering, contribute upstream (open-source/operators), and participate in a 24/7 on-call rotation to resolve deployment infrastructure issues.
Top Skills: AlertingArgo WorkflowsArgocdAWSAzureCi/CdContainersDnsGCPGoKubernetesLinuxLoad BalancerObservabilityPythonService MeshTcp/IpTls
Reposted 3 Days AgoSaved
Remote
New York, NY
Mid level
Mid level
Blockchain • Software
Build, operate, and scale production Kubernetes infrastructure using GitOps and declarative IaC. Design CI/CD workflows, observability, and secure-by-default systems. Troubleshoot networking/storage, participate in on-call rotations, automate operational workflows, and drive postmortems and reliability improvements.
Top Skills: ArbitrumArgocdArgocd ApplicationsetsAWSAzureBashCloudwatchCodebuildGCPGithub ActionsGitopsGoGrafanaK9SKubernetesLinuxLokiMimirPrometheusPrysmPythonTerraformYamlZerodev
3 Days AgoSaved
Remote
New York, NY
177K-240K Annually
Expert/Leader
177K-240K Annually
Expert/Leader
Information Technology • Security • Software • Cybersecurity
Lead reliability strategy across eight engineering groups by defining SLIs, SLOs, and error budgets; strengthening incident response, alerting, postmortems, change safety, and failure testing; and coaching teams to own reliability. The role remains hands-on through production investigations, tooling, dashboards, and reference implementations. It also leads AI adoption in incident management and observability while partnering with architecture, platform, and product teams to improve distributed-system resilience.
Top Skills: Ai ToolsAWSDatadogDynamoDBElasticsearchGoInfrastructure As CodeKafkaKubernetesObservability ToolingRedisTypescript
13 Days AgoSaved
Hybrid
New York, NY
183K-247K Annually
Senior level
183K-247K Annually
Senior level
Edtech • Machine Learning • Mobile • Other • Software
Designs, operates, and improves Duolingo’s large-scale distributed systems and core infrastructure. Responsibilities include diagnosing production issues, developing platforms and automation, conducting launch reviews and root cause analyses, maintaining incident response and postmortem practices, and improving reliability, scalability, and engineering velocity. The role partners with product and platform engineering teams to reduce operational toil and ensure high-quality service delivery.
Top Skills: DockerDynamoGoJavaKotlinKubernetesMesosMySQLNomadPostgresPython
13 Days AgoSaved
In-Office or Remote
New York, NY
320K-485K Annually
Senior level
320K-485K Annually
Senior level
Artificial Intelligence • Natural Language Processing • Generative AI
Own production safeguards infrastructure for Claude model launches and safety classifier deployments. Configure and verify safeguards across first-party, AWS Bedrock, and GCP Vertex platforms; lead canary rollouts, post-deployment validation, incident response, and rollback decisions. Build automation, continuous validation, repeatable deployment pipelines, and a provenance registry to reduce operational toil and configuration drift. Participate in on-call rotations and launch readiness processes.
Top Skills: AWSGCPLlm Inference SystemsMl InfrastructurePythonRustTransformer-Based Models
Reposted 26 Days AgoSaved
In-Office or Remote
New York, NY
Mid level
Mid level
Renewable Energy
Own reliability, performance, and scalability of Postgres and ClickHouse databases. Build scalable data pipelines, design analytical schemas and DBT models, migrate data to ClickHouse, implement data quality checks, eliminate duplicates, and manage database infrastructure via IaC.
Top Skills: Aws CdkAws Step FunctionsCi/CdClickhouseDagsterDbtPostgresPulumiPythonSQL
Reposted 26 Days AgoSaved
Remote
New York, NY
Senior level
Senior level
Software
As a Senior DevOps / Platform Reliability Engineer, you will manage CI/CD pipelines, automate infrastructure, operate Kubernetes, and enhance observability while ensuring security and compliance for enterprise systems.
Top Skills: Argo CdAurora MysqlAWSBashCloudFormationEksElasticacheGithub ActionsGrafanaKubernetesLinuxMskOpentelemetryPrometheusPythonS3Terraform
Reposted One Month AgoSaved
In-Office
New York, NY
173K-224K Annually
Mid level
173K-224K Annually
Mid level
Artificial Intelligence • Software
Own reliability for named customer workloads; debug distributed systems across hardware, fabric, and scheduler; run customer-facing incident communications; convert recurring customer pain into engineering fixes; support large-scale compute customers and push internal teams to resolve root causes.
Top Skills: CloudGpuHpcInfinibandKubernetesNcclRoceSlurm
One Month AgoSaved
In-Office or Remote
New York, NY
145K-175K Annually
Senior level
145K-175K Annually
Senior level
Information Technology • Business Intelligence • Consulting
Lead hands-on SRE/DevOps practitioner who drives safe, automated deployments and CI/CD best practices. Wrap applications with test suites, containerize, upgrade runtimes and dependencies, improve code review and security practices, implement observability, automate deployments (Ansible, Liquibase), and enable resilient cross-datacenter topologies.
Top Skills: AnsibleDockerGradleJavaJvm LanguagesLiquibaseMavenNpmOpenshiftRedhat LinuxRenovateSpring FrameworkTest-Driven Development (Tdd)Typescript
Reposted One Month AgoSaved
Easy Apply
Remote
New York, NY
Easy Apply
100K-110K Annually
Mid level
100K-110K Annually
Mid level
Healthtech • Software
Operate and maintain AWS-hosted MERN applications and large-scale data workflows. Manage serverless and Spark-based pipelines, perform incident response and on-call duties, engineer automation to eliminate operational toil, ensure HIPAA/SOC2/HITRUST compliance, build observability and lead blameless post-mortems.
Top Skills: Amazon EcsAmazon EksAmazon EmrAthenaAws GlueAws LambdaAws SnsAws SqsCloudwatchEc2IamJavaScriptMernMySQLNode.jsOpentofuPysparkPythonRabbitMQTerraformTypescriptVpc
One Month AgoSaved
Remote or Hybrid
New York, NY
140K-215K Annually
Senior level
140K-215K Annually
Senior level
Cloud • Computer Vision • Information Technology • Sales • Security • Cybersecurity
Senior SRE owning availability, automation, and observability for CI/CD platform services. Build and operate infrastructure, run on-call, lead incident response, mentor engineers, drive design/capacity planning, integrate AI-assisted workflows, and improve cross-team reliability.
Top Skills: Active DirectoryAnsibleApache AirflowSparkAWSAzureBashBazelBitbucketCassandraChefDatadogDnsFirewall RulesGCPGitGithub ActionsGitlabGitlab CiGoGrafanaHoneycombHumio/LogscaleJenkinsKafkaKubernetesLoad BalancersMongoDBMySQLNasNew RelicNfsObject StorageOpensearchOraclePostgresPowershellPrometheusPulsarPuppetPythonRabbitMQRedis/ValkeyRedpandaRoutingSaltSanSplunkTerraformVarnishVipsWindows Server
Reposted 28 Days AgoSaved
In-Office or Remote
New York, NY
119K-178K Annually
Senior level
119K-178K Annually
Senior level
Automotive • Information Technology • Other • Transportation • Energy
Perform RAM and FMECA/FMEA analyses, develop fault trees and reliability predictions, support maintainability and logistics analyses, produce reliability growth test plans, contribute to systems engineering documentation, advise design engineers on R&M shortfalls, and present results to management and clients.
Top Skills: Fault Tree AnalysisFmeaFmecaIntegrated Logistics Support (Ils/Ilsa)Iso-9000Mil-Hdbk-217FRam ModellingRam SoftwareStatistical Methods
Reposted 28 Days AgoSaved
In-Office or Remote
New York, NY
Senior level
Senior level
Artificial Intelligence • Cloud • Information Technology • Software
The Site Reliability Engineer will provision and manage Kubernetes clusters, build automation tools, debug customer issues, and improve infrastructure reliability.
Top Skills: AnsibleBashDatadogGoGrafanaHelmKubernetesLokiPrometheusPythonTerraform
All Filters
JobType
New Jobs
Job Category
Experience
Industry
Company Name
Company Size

Sign up now Access later

Create Free Account