Maximum of 25 job preferences reached.
Top Reliability Engineer Jobs in NYC, NY
Artificial Intelligence • Information Technology
Own the reliability, performance, and uptime of large-scale post-training and reinforcement learning systems. Debug distributed failures across accelerators, networking, storage, schedulers, and training frameworks; build monitoring, alerting, recovery, checkpointing, and scheduling tools; improve cluster utilization and fault tolerance; support production model runs through on-call rotations and postmortems; and partner closely with research teams during active training runs.
Top Skills:
C++DpoGoGpuInfinibandKubernetesLinuxNcclPpoPythonPyTorchRayRdmaRlhfSlurmTpu
Artificial Intelligence • Information Technology
Owns end-to-end reliability for Tinker, including CI/CD, observability, service-level objectives, incident response, multi-tenant isolation, resource scheduling, and vulnerability remediation. The role designs monitoring for distributed training systems, improves recovery and checkpointing, and operates Kubernetes clusters supporting heterogeneous GPU workloads. It collaborates closely with platform, security, engineering, and research teams to improve production resilience and utilization.
Top Skills:
Ci/CdCloud InfrastructureDistributed SystemsGpu WorkloadsKubernetesLoraProduction Observability
Artificial Intelligence • HR Tech • Professional Services
Maintain and support platform environments across the software development lifecycle. Configure monitoring, alerting, and self-healing solutions; troubleshoot application and infrastructure issues with developers and testers; maintain technical documentation; design scalable architectural solutions; and mentor junior engineers. The role requires cloud, Kubernetes, operating system, database, scripting, configuration management, .NET, and AI/ML automation expertise.
Top Skills:
.NetAi/MlAnsibleAWSAzureBashGCPHelmJenkinsKubernetesLinuxOracle CloudOracle DatabasePowershellService MeshSQL ServerWindows
Cloud • Security • Software • Cybersecurity
Architect, develop, test, and distribute software, services, and infrastructure supporting Akamai’s cloud hypervisor platforms. Improve observability, automate infrastructure processes, troubleshoot complex distributed-system issues, mentor engineers, and participate in on-call service restoration. The role requires deep Linux, kernel, virtualization, ARM hardware, large-scale infrastructure, DevOps, and configuration-management expertise.
Top Skills:
AnsibleArmDevOpsDistributed SystemsKvm/QemuLinuxLinux KernelNested VirtualizationNvidia GraceObservability InfrastructureSaltstack
Cloud • Security • Software • Cybersecurity
Architect, build, and support reliable network infrastructure and automation for Akamai’s distributed cloud platform. Develop Bash and Python tooling, establish deployment standards, define SLOs, mentor engineers, and troubleshoot complex network issues. The role requires expertise in large-scale distributed systems, TCP/IP, BGP, Linux networking, configuration management, CI/CD, and open-source networking software, with participation in on-call rotations.
Top Skills:
AnsibleArgocdBashBgpBirdChefFirewallsFrrGithub ActionsGoGobgpJenkinsLinux NetworkingLoad BalancingPuppetPythonRustSaltstackSlack BotsTcp/Ip
Software
Own reliability, performance, scalability, and operational standards across on-premises, private-cloud, and AWS environments. Build infrastructure-as-code, deployment automation, monitoring, alerting, and observability tooling; define SLOs, SLIs, error budgets, and readiness standards. Lead incident response, root-cause analysis, disaster-recovery readiness, release coordination, and preventive automation. Mentor engineers, coach teams on operational practices, coordinate on-call coverage, and ensure infrastructure meets security and compliance requirements.
Top Skills:
AnsibleAuto ScalingAWSBashCi/CdClaudeCloudwatchCrowdstrikeDatadogDistributed SystemsDnsDockerEc2EcsGitGithub CopilotGitlabGitopsIamKubernetesLinuxLoad BalancingOracle LinuxPythonQualysRapid7RhelS3Tcp/IpTerraformVpcWireshark
Healthtech • Software
Design, automate, and maintain scalable infrastructure and SRE tooling. Manage Kubernetes clusters, CI/CD, monitoring, and incident response. Improve processes, reduce toil via automation, and collaborate with engineering and data teams to support domestic and international workloads.
Top Skills:
AWSAzureContainerdDnsDockerFirewallsGCPGoGrpcHelmKubernetesLinuxLoad BalancingPrometheusPythonRoutingShell ScriptingTcp/IpUdp
Cloud • Software
The Senior Site Reliability Engineer improves production reliability, resilience, and system availability. Responsibilities include operating Linux and Kubernetes infrastructure, automating workflows, developing tools, managing observability, responding to incidents, participating in on-call rotations, writing runbooks, and conducting blameless postmortems. The role collaborates with distributed teams and stakeholders, supports cloud infrastructure, troubleshoots live systems, and promotes DevOps and SRE best practices.
Top Skills:
AWSBare-Metal InfrastructureBashCassandraCi/CdClickhouseGitGoKubernetesLinuxLogsMetricsObservabilityPostgresPythonTerraformTraces
Artificial Intelligence • Marketing Tech • Software
Lead SRE to provide technical leadership for a multi-cloud, multi-region active-active content platform. Define automation and IaC strategy, design core platform and logging/observability systems, establish capacity and performance frameworks, lead cross-team reliability initiatives, improve on-call and incident response, mentor engineers, and proactively address systemic reliability and scaling challenges.
Top Skills:
Apache KafkaApache PulsarAWSCassandraChefEksGCPGkeGoGrafana AlloyGrafana LokiKubernetesLinuxNode.jsPrometheusPythonRubyScylladbShell ScriptingTempoTerraformThanos
Software
Owns reliability, observability, performance, and security for a multi-region SaaS platform. Responsibilities include managing Datadog, implementing APM and tracing, defining SLOs, developing automation, expanding infrastructure as code and CI/CD, automating operational workflows, maintaining security controls, participating in incident response, documenting procedures, and mentoring engineers.
Top Skills:
ApmAzure DevopsAzure Kubernetes ServiceAzure SqlBashBicepCi/CdCosmos DbDatadogDistributed TracingHelmInfrastructure As CodeKey VaultKubernetesKustomizeManaged IdentitiesAzureMicrosoft Entra IdPowershellPythonRedisService BusTerraform
Artificial Intelligence • Hardware • Software • Semiconductor
Develop kernel-centric reliability solutions for AI compute clusters and production services. Responsibilities include debugging failures, building diagnostic tools, supporting incident response, performing root-cause analysis, improving kernel and software reliability, and collaborating with systems, hardware, ASIC, and architecture teams on reliability-focused designs.
Top Skills:
CC++Core Dump HandlingDebuggersDistributed ProgrammingGpusParallel ProgrammingProfilersPythonSanitizersTracing
Healthtech • Pharmaceutical • Manufacturing
Support operational deployments and maintenance for Core Speech production systems. Implement and maintain monitoring, alerting, performance reporting, and capacity planning. Participate in on-call rotation for off-hours reliability support. Drive improvements to system performance and architecture, and collaborate on projects while complying with corporate and quality standards.
New
Track Smarter, Apply Better.
Ditch the spreadsheets. Organize your job search with our freeApplication Tracker.
Use For Free
Reposted One Month AgoSaved
Easy Apply
Easy Apply
Big Data • Cloud • Software • Database
Maintain and improve multi-cloud Kubernetes infrastructure, CI/CD (Argo Workflows/ArgoCD), observability, and networking. Build reliable continuous deployment tooling and onboarding flows, provide internal support, collaborate across Platform Engineering, contribute upstream (open-source/operators), and participate in a 24/7 on-call rotation to resolve deployment infrastructure issues.
Top Skills:
AlertingArgo WorkflowsArgocdAWSAzureCi/CdContainersDnsGCPGoKubernetesLinuxLoad BalancerObservabilityPythonService MeshTcp/IpTls
Blockchain • Software
Build, operate, and scale production Kubernetes infrastructure using GitOps and declarative IaC. Design CI/CD workflows, observability, and secure-by-default systems. Troubleshoot networking/storage, participate in on-call rotations, automate operational workflows, and drive postmortems and reliability improvements.
Top Skills:
ArbitrumArgocdArgocd ApplicationsetsAWSAzureBashCloudwatchCodebuildGCPGithub ActionsGitopsGoGrafanaK9SKubernetesLinuxLokiMimirPrometheusPrysmPythonTerraformYamlZerodev
Information Technology • Security • Software • Cybersecurity
Lead reliability strategy across eight engineering groups by defining SLIs, SLOs, and error budgets; strengthening incident response, alerting, postmortems, change safety, and failure testing; and coaching teams to own reliability. The role remains hands-on through production investigations, tooling, dashboards, and reference implementations. It also leads AI adoption in incident management and observability while partnering with architecture, platform, and product teams to improve distributed-system resilience.
Top Skills:
Ai ToolsAWSDatadogDynamoDBElasticsearchGoInfrastructure As CodeKafkaKubernetesObservability ToolingRedisTypescript
Edtech • Machine Learning • Mobile • Other • Software
Designs, operates, and improves Duolingo’s large-scale distributed systems and core infrastructure. Responsibilities include diagnosing production issues, developing platforms and automation, conducting launch reviews and root cause analyses, maintaining incident response and postmortem practices, and improving reliability, scalability, and engineering velocity. The role partners with product and platform engineering teams to reduce operational toil and ensure high-quality service delivery.
Top Skills:
DockerDynamoGoJavaKotlinKubernetesMesosMySQLNomadPostgresPython
Artificial Intelligence • Natural Language Processing • Generative AI
Own production safeguards infrastructure for Claude model launches and safety classifier deployments. Configure and verify safeguards across first-party, AWS Bedrock, and GCP Vertex platforms; lead canary rollouts, post-deployment validation, incident response, and rollback decisions. Build automation, continuous validation, repeatable deployment pipelines, and a provenance registry to reduce operational toil and configuration drift. Participate in on-call rotations and launch readiness processes.
Top Skills:
AWSGCPLlm Inference SystemsMl InfrastructurePythonRustTransformer-Based Models
Renewable Energy
Own reliability, performance, and scalability of Postgres and ClickHouse databases. Build scalable data pipelines, design analytical schemas and DBT models, migrate data to ClickHouse, implement data quality checks, eliminate duplicates, and manage database infrastructure via IaC.
Top Skills:
Aws CdkAws Step FunctionsCi/CdClickhouseDagsterDbtPostgresPulumiPythonSQL
Software
As a Senior DevOps / Platform Reliability Engineer, you will manage CI/CD pipelines, automate infrastructure, operate Kubernetes, and enhance observability while ensuring security and compliance for enterprise systems.
Top Skills:
Argo CdAurora MysqlAWSBashCloudFormationEksElasticacheGithub ActionsGrafanaKubernetesLinuxMskOpentelemetryPrometheusPythonS3Terraform
Artificial Intelligence • Software
Own reliability for named customer workloads; debug distributed systems across hardware, fabric, and scheduler; run customer-facing incident communications; convert recurring customer pain into engineering fixes; support large-scale compute customers and push internal teams to resolve root causes.
Top Skills:
CloudGpuHpcInfinibandKubernetesNcclRoceSlurm
Information Technology • Business Intelligence • Consulting
Lead hands-on SRE/DevOps practitioner who drives safe, automated deployments and CI/CD best practices. Wrap applications with test suites, containerize, upgrade runtimes and dependencies, improve code review and security practices, implement observability, automate deployments (Ansible, Liquibase), and enable resilient cross-datacenter topologies.
Top Skills:
AnsibleDockerGradleJavaJvm LanguagesLiquibaseMavenNpmOpenshiftRedhat LinuxRenovateSpring FrameworkTest-Driven Development (Tdd)Typescript
Healthtech • Software
Operate and maintain AWS-hosted MERN applications and large-scale data workflows. Manage serverless and Spark-based pipelines, perform incident response and on-call duties, engineer automation to eliminate operational toil, ensure HIPAA/SOC2/HITRUST compliance, build observability and lead blameless post-mortems.
Top Skills:
Amazon EcsAmazon EksAmazon EmrAthenaAws GlueAws LambdaAws SnsAws SqsCloudwatchEc2IamJavaScriptMernMySQLNode.jsOpentofuPysparkPythonRabbitMQTerraformTypescriptVpc
Cloud • Computer Vision • Information Technology • Sales • Security • Cybersecurity
Senior SRE owning availability, automation, and observability for CI/CD platform services. Build and operate infrastructure, run on-call, lead incident response, mentor engineers, drive design/capacity planning, integrate AI-assisted workflows, and improve cross-team reliability.
Top Skills:
Active DirectoryAnsibleApache AirflowSparkAWSAzureBashBazelBitbucketCassandraChefDatadogDnsFirewall RulesGCPGitGithub ActionsGitlabGitlab CiGoGrafanaHoneycombHumio/LogscaleJenkinsKafkaKubernetesLoad BalancersMongoDBMySQLNasNew RelicNfsObject StorageOpensearchOraclePostgresPowershellPrometheusPulsarPuppetPythonRabbitMQRedis/ValkeyRedpandaRoutingSaltSanSplunkTerraformVarnishVipsWindows Server
Automotive • Information Technology • Other • Transportation • Energy
Perform RAM and FMECA/FMEA analyses, develop fault trees and reliability predictions, support maintainability and logistics analyses, produce reliability growth test plans, contribute to systems engineering documentation, advise design engineers on R&M shortfalls, and present results to management and clients.
Top Skills:
Fault Tree AnalysisFmeaFmecaIntegrated Logistics Support (Ils/Ilsa)Iso-9000Mil-Hdbk-217FRam ModellingRam SoftwareStatistical Methods
Artificial Intelligence • Cloud • Information Technology • Software
The Site Reliability Engineer will provision and manage Kubernetes clusters, build automation tools, debug customer issues, and improve infrastructure reliability.
Top Skills:
AnsibleBashDatadogGoGrafanaHelmKubernetesLokiPrometheusPythonTerraform
Let Your Resume Do The Work
Upload your resume to be matched with jobs you're a great fit for.
Success! We'll use this to further personalize your experience.
Top NYC, NY Companies Hiring Reliability Engineers
See AllPopular NYC, NY Engineering Job Searches
Engineering Jobs in NYC, NY
.NET Developer Jobs in NYC, NY
Android Developer Jobs in NYC, NY
Associate Software Engineer Jobs in NYC, NY
Automation Engineer Jobs in NYC, NY
AWS Engineer Jobs in NYC, NY
Backend Engineer Jobs in NYC, NY
C# Jobs in NYC, NY
C++ Jobs in NYC, NY
Cloud Engineer Jobs in NYC, NY
Controls Engineer Jobs in NYC, NY
CTO Jobs in NYC, NY
Design Engineer Jobs in NYC, NY
DevOps Engineer Jobs in NYC, NY
DevOps Jobs in NYC, NY
Director of Engineering Jobs in NYC, NY
Electrical Engineering Jobs in NYC, NY
Embedded Software Engineer Jobs in NYC, NY
Engineering Manager Jobs in NYC, NY
Field Engineer Jobs in NYC, NY
Front End Developer Jobs in NYC, NY
Full-Stack Engineer Jobs in NYC, NY
Golang Jobs in NYC, NY
Hardware Engineer Jobs in NYC, NY
Infrastructure Engineer Jobs in NYC, NY
iOS Developer Jobs in NYC, NY
Java Developer Jobs in NYC, NY
Java Full-Stack Engineer Jobs in NYC, NY
Javascript Jobs in NYC, NY
Lead Software Engineer Jobs in NYC, NY
Linux Jobs in NYC, NY
Manufacturing Engineer Jobs in NYC, NY
Mechanical Design Engineer Jobs in NYC, NY
Mechanical Engineering Jobs in NYC, NY
Mechatronics Engineering Jobs in NYC, NY
Network Engineer Jobs in NYC, NY
PHP Developer Jobs in NYC, NY
Platform Engineer Jobs in NYC, NY
Principal Architect Jobs in NYC, NY
Principal Engineer Jobs in NYC, NY
Principal Software Engineer Jobs in NYC, NY
Process Engineer Jobs in NYC, NY
Project Engineer Jobs in NYC, NY
Python Jobs in NYC, NY
QA Engineer Jobs in NYC, NY
QA Jobs in NYC, NY
Quality Assurance Automation Engineer Jobs in NYC, NY
Reliability Engineer Jobs in NYC, NY
Robotics Engineer Jobs in NYC, NY
Ruby Jobs in NYC, NY
Salesforce Developer Jobs in NYC, NY
Scala Jobs in NYC, NY
Security Engineer Jobs in NYC, NY
Software Engineer Jobs in NYC, NY
Solutions Architect Jobs in NYC, NY
Solutions Engineer Jobs in NYC, NY
SRE Engineer Jobs in NYC, NY
Staff Engineer Jobs in NYC, NY
Staff Software Engineer Jobs in NYC, NY
Systems Engineer Jobs in NYC, NY
Vice President Of Engineering Jobs in NYC, NY
Web Developer Jobs in NYC, NY
All Filters
Total selected ()
No Results
No Results






























