Maximum of 25 job preferences reached.
Top SRE Engineer Jobs in NYC, NY
AdTech • Big Data • Internet of Things • Marketing Tech • Mobile • Software • Analytics
Senior Site Reliability Engineer responsible for improving platform availability, scalability, resilience, and observability. The role manages AWS and EKS infrastructure with Terraform, builds CI/CD pipelines using GitHub Actions, expands GitOps through ArgoCD and Helm, strengthens disaster recovery, supports incident response, and develops monitoring, alerting, dashboards, and SLOs. The engineer partners with product teams, leads infrastructure initiatives, troubleshoots Kubernetes, and supports high-availability distributed systems.
Top Skills:
Amazon EksAmazon RdsAmazon VpcArgocdAWSAws IamCluster AutoscalerCrossplaneDatadogGithub ActionsGoHelmKarpenterKubernetesOpentofuPrometheusPythonShell ScriptingTerraform
Artificial Intelligence • eCommerce • Retail • Software
Build and maintain CI/CD pipelines, manage and automate cloud infrastructure and configurations, implement monitoring/logging and alerting for reliability, enforce security and compliance practices, and collaborate with development teams to support scaling and operations.
Top Skills:
Soc ISoc Ii
Cloud • Healthtech • Internet of Things • Machine Learning • Software
Leads reliability engineering for a cloud-based, real-time healthcare platform. Responsibilities include maintaining 99.99% uptime, designing fault-tolerant and self-healing systems, managing AWS and Kubernetes infrastructure, implementing Terraform and observability tooling, leading incident response and disaster recovery, ensuring HIPAA and SOC 2 compliance, and driving scalability strategy while mentoring the SRE team.
Top Skills:
AWSCi/CdDatadogDistributed SystemsDockerGitopsGrafanaKubernetesOpentelemetryPrometheusTerraform
Cloud
Design, automate, and maintain highly available cloud infrastructure and CI/CD platforms. Troubleshoot and debug Linux systems and networking, automate infrastructure (Terraform/Chef/Ansible/Puppet), and improve scalability, reliability, and platform velocity. Participate in on-call rotation and collaborate across engineering teams.
Top Skills:
AnsibleApache HttpdApache TomcatAWSBashChefCi/CdDnsDockerFedrampGdbGitGoHTTPKubernetesLinuxLoad BalancingLtraceNginxPki/Federated Certificate ManagementPuppetPythonStraceTcp/IpTcpdumpTerraformWireshark
Edtech • Kids + Family • Sports
Audit infrastructure, deployment pipelines, monitoring, alerting, incident response, on-call practices, and internal tools. Produce actionable audit reports, implement code and configuration fixes, improve SLOs and reliability practices, advise on scalable architecture, and partner with engineers on implementation and handoff. The role is a fully remote, part-time consulting engagement with potential for full-time conversion.
Top Skills:
Ai Coding ToolsAWSCi/CdDatadogGrafanaPrometheus
Cloud • Information Technology • Cybersecurity • Infrastructure as a Service (IaaS)
Owns reliability, observability, and incident response for a GPUaaS platform. Defines SLOs, builds monitoring and alerting systems, leads major incidents and post-incident reviews, automates operational processes, maintains runbooks, manages on-call operations, coordinates with engineering teams, drives chaos testing, reports SLA performance, and mentors junior engineers.
Top Skills:
DatadogGoGpuaasGrafanaGremlinHpcKubernetesLitmusOpentelemetryPrometheusPython
Cloud
The Senior Site Reliability Engineer will enhance the Splunk ecosystem and develop an Observability Platform by automating infrastructure and managing complex distributed systems, while optimizing log collection and incident response.
Top Skills:
AWSGCPGoKubernetesLinuxOpentelemetryPythonRubySplunkTerraform
Cloud • Healthtech • Internet of Things • Machine Learning • Software
Lead reliability, performance, automation, and scalability for a 24/7 AWS-based healthcare platform. Responsibilities include designing fault-tolerant infrastructure, managing Kubernetes and Docker environments, implementing Terraform and GitOps, improving observability, leading incident response and disaster recovery, and ensuring HIPAA and SOC 2 compliance. The role also sets technical strategy, reduces downtime through self-healing automation, and leads and mentors the SRE team.
Top Skills:
AWSCi/CdDatadogDockerFhirGitopsGrafanaHipaaHl7KubernetesOpentelemetryPrometheusSoc 2Terraform
Artificial Intelligence • Cloud • Fintech • Machine Learning • Mobile • Software
Lead design, development, deployment, and scaling of cloud infrastructure and SRE tooling. Build automation, CI/CD, observability, capacity planning, and reliability improvements; collaborate with product teams to define non-functional requirements and resolve production issues.
Top Skills:
.NetApi GatewayAWSAzureC#Data LakehouseDatabricks DeltaDatadogElasticsearchElkEvent HubsFunctions/ServerlessGitGrafanaJavaJenkinsKafkaKibanaKubernetesLogstashPowershellSnowflakeSqsTeamcityVisual Basic
Healthtech • Software
Manage and optimize a multi-account AWS environment supporting production healthcare applications and analytics platforms. Responsibilities include AWS infrastructure administration, CI/CD and infrastructure automation, observability, incident response, on-call support, disaster recovery validation, HIPAA/HiTrust compliance, IAM and security controls, and support for containerized Java and Python applications. The role also contributes to Kubernetes and EKS modernization initiatives and partners with developers to resolve complex production issues.
Top Skills:
Amazon EksApi GatewayArgocdAuroraAWSAws CloudformationAws CodepipelineAws Security HubBashCloudfrontCloudwatchCortex CloudDatadogDockerEc2EcsFargateGitGuarddutyHelmIamJavaJenkinsKubernetesLambdaLinuxPrismaPythonRdsS3Spring BootUbuntuZabbix
Software
Owns reliability, observability, performance, and security for a multi-region SaaS platform. Responsibilities include managing Datadog, implementing APM and tracing, defining SLOs, developing automation, expanding infrastructure as code and CI/CD, automating operational workflows, maintaining security controls, participating in incident response, documenting procedures, and mentoring engineers.
Top Skills:
ApmAzure DevopsAzure Kubernetes ServiceAzure SqlBashBicepCi/CdCosmos DbDatadogDistributed TracingHelmInfrastructure As CodeKey VaultKubernetesKustomizeManaged IdentitiesAzureMicrosoft Entra IdPowershellPythonRedisService BusTerraform
Fintech • Payments • Financial Services
Operate and maintain AWS infrastructure and CI/CD pipelines, deploy containerized applications, write Terraform IaC, implement monitoring/observability, troubleshoot incidents, resolve security vulnerabilities, participate in on-call rotation, and produce operational/runbook documentation to ensure resilient cloud platform delivery.
Top Skills:
AgileAlbAmazon Web Services (Aws)Aws IamAws LambdaCloudbeaverCloudwatchDevsecopsDirect ConnectDnsEc2EcsEfsEksElbGitlab CiGlueGrafanaHelmJavaJbossJIRANginxOktaOracle RdsPythonRds PostgresRoute53S3ScalaSnsTerraformTomcatTransit GatewayVpc
New
Cut your apply time in half.
Use ourAI Assistantto automatically fill your job applications.
Use For Free
Artificial Intelligence • Cloud • Software • Infrastructure as a Service (IaaS)
Ensure stability and resilience of Runpod's distributed AI platform by defining SLIs/SLOs, leading incident response, building observability and reliability tooling, automating operational workflows, and partnering with engineering teams to reduce toil and improve production readiness.
Top Skills:
BashCi/CdContainerized Production SystemsGoGpu Observability ToolingGrafanaInfrastructure As CodeLinuxPrometheusPython
Edtech
Lead infrastructure modernization and platform reliability across multiple cloud providers. Design infrastructure as code, operate Kubernetes and Linux environments, improve CI/CD and deployment tooling, establish SLI/SLO practices, strengthen observability, lead incident response, manage cloud costs, and partner on security and compliance. Provide technical leadership through architecture guidance, mentorship, engineering standards, and roadmap development while participating in on-call support.
Top Skills:
AWSCi/CdGCPJenkinsKubernetesLinuxPythonRubyRuby On RailsSoc 2SpinnakerTerraform
Cloud • Information Technology • Business Intelligence • Consulting
Design, build, and operate cloud infrastructure and SRE capabilities for an enterprise AI platform. Responsibilities include infrastructure-as-code, landing zones, networking, Kubernetes, CI/CD, observability, incident response, SLOs, production readiness, automation, cost optimization, and support for hybrid, edge, air-gapped, and customer-controlled environments. The role is remote, client-facing, and requires strong collaboration and reliability ownership.
Top Skills:
AlertingAzureAzure ArcAzure DevopsBicepCi/CdDashboardsDockerGithub ActionsGpu WorkloadsInfrastructure As CodeKubernetesLogsMetricsObservabilitySlosTerraformTraces
Cloud • Security • Software • Cybersecurity
The Site Reliability Engineer II ensures the reliability, availability, performance, and security of critical cloud systems and services. Responsibilities include developing automation for provisioning and configuration management, maintaining monitoring and alerting, optimizing infrastructure performance, supporting high availability, and enabling continuous integration and delivery. The role collaborates with security teams and drives operational improvements across cloud and network infrastructure.
Top Skills:
AnsibleAWSAzureChefContinuous DeliveryContinuous IntegrationDnsElk StackGCPGoGrafanaHTTPKubernetesLinuxPrometheusPuppetPythonShellTcp/IpUnix
Fintech • Software • Financial Services
Own AWS and Kubernetes reliability, deployments, observability, incident response, infrastructure as code, CI/CD, and compliance controls for SOC 2 and PCI-DSS. Build scalable platform foundations, AI infrastructure, and greenfield product infrastructure while improving cost efficiency, security, vendor flexibility, and deployment consistency. Participate in on-call support and partner with security and engineering teams across backend, mobile, data, ML, and AI.
Top Skills:
Ai/Llm ToolingAmazon EksAWSAws CdkCi/CdDatadogKotlinKubernetesMcpNode.jsOpentelemetryOpentofuPci-DssPythonSoc 2TerraformTypescript
Reposted One Month AgoSaved
Easy Apply
Easy Apply
Cloud • Information Technology • Security • Software • Cybersecurity
This internship role focuses on SRE skills, requiring collaboration and problem-solving in dynamic environments for Zscaler's Zero Trust Exchange team.
Top Skills:
AnsibleAws EcsKubernetesLinuxPythonTerraform
Cloud • Security • Software • Cybersecurity
Build and maintain reliable, scalable cloud compute platforms across distributed services. Troubleshoot Linux, networking, and production issues; develop automation and AI-assisted tooling; improve monitoring, alerting, SLIs, and SLOs; conduct incident response and root cause analysis; and partner with engineering teams on system design, deployment safety, and operational readiness.
Top Skills:
AnsibleDnsDockerElkGoGrafanaKubernetesLinuxLokiNomadOpensearchPodmanPrometheusPythonSaltTcp/IpTerraform
Payments • Software • Automation
Lead platform and infrastructure direction on AWS, evolve CI/CD and ephemeral environments, set observability and SLO standards, drive incident response and postmortems, mentor engineers, and build automation to reduce operational risk.
Top Skills:
AWSCi/CdDistributed SystemsEcsEphemeral Environments/Preview DeploysFargateGithub ActionsLogsObservability (MetricsSlos/Slis/Error BudgetsTracing)
Software • Cybersecurity
This role involves managing Kubernetes clusters, cloud infrastructure, and CI/CD pipelines. The engineer will enhance system reliability and efficiency while troubleshooting production issues.
Top Skills:
AlertmanagerAWSAzureBashCi/CdDockerElastic StackElasticsearchGCPGoGrafanaHelmKafkaKubernetesLokiMongoDBOciPrometheusPythonRedisSparkTerraform
Fintech • Payments • Financial Services
The Site Reliability Engineer will assist clients with Redline products, manage production environments, troubleshoot issues, and ensure automation and customer satisfaction.
Top Skills:
C/C++JavaLinuxPython
Artificial Intelligence • Cloud • Social Impact • Software • Wearables
Own reliability, observability, and operational excellence for the Unified Call (911) platform. Design monitoring, alerting, incident response, deployment automation, and dashboards. Improve system resiliency, analyze architecture for operational risks, and build tooling and practices to enable reliable production operations across a Kubernetes-based cloud environment.
Top Skills:
AWSDatadogKafkaKubernetesRabbitMQ
Artificial Intelligence • Healthtech • Machine Learning • Software
Define reliability strategy and technical roadmaps; design scalable cloud infrastructure; improve observability, incident response, disaster recovery, automation, and deployment systems. Partner across engineering, security, data, AI, and product teams to strengthen resilience, compliance, and operational excellence. Lead architecture reviews, resolve complex production issues, establish SRE practices, and mentor engineers while reducing operational toil and improving platform scalability.
Top Skills:
AutomationCi/CdCloud InfrastructureContainer OrchestrationContainersDisaster RecoveryDistributed SystemsInfrastructure As CodeMulti-Region ArchitectureNetworkingObservability
Natural Language Processing • Software • Conversational AI
Design, deploy, and manage resilient GCP network architectures, including VPCs, Interconnects, load balancing, and DNS. Lead Terraform-based infrastructure automation, implement network security policies, optimize cloud networking costs, and resolve connectivity and performance issues. Drive SRE practices such as capacity planning, error budgets, monitoring, and alerting. Partner with SRE, platform, and engineering teams to onboard services and maintain high availability, reliability, and security across distributed systems.
Top Skills:
BashCloud InterconnectDnsGoGoogle Cloud Platform (Gcp)Infrastructure As Code (Iac)LinuxLoad BalancingPythonTerraformVpc
Let Your Resume Do The Work
Upload your resume to be matched with jobs you're a great fit for.
Success! We'll use this to further personalize your experience.
Top NYC, NY Companies Hiring SRE Engineers
See AllPopular NYC, NY Engineering Job Searches
Engineering Jobs in NYC, NY
.NET Developer Jobs in NYC, NY
Android Developer Jobs in NYC, NY
Associate Software Engineer Jobs in NYC, NY
Automation Engineer Jobs in NYC, NY
AWS Engineer Jobs in NYC, NY
Backend Engineer Jobs in NYC, NY
C# Jobs in NYC, NY
C++ Jobs in NYC, NY
Cloud Engineer Jobs in NYC, NY
Controls Engineer Jobs in NYC, NY
CTO Jobs in NYC, NY
Design Engineer Jobs in NYC, NY
DevOps Engineer Jobs in NYC, NY
DevOps Jobs in NYC, NY
Director of Engineering Jobs in NYC, NY
Electrical Engineering Jobs in NYC, NY
Embedded Software Engineer Jobs in NYC, NY
Engineering Manager Jobs in NYC, NY
Field Engineer Jobs in NYC, NY
Front End Developer Jobs in NYC, NY
Full-Stack Engineer Jobs in NYC, NY
Golang Jobs in NYC, NY
Hardware Engineer Jobs in NYC, NY
Infrastructure Engineer Jobs in NYC, NY
iOS Developer Jobs in NYC, NY
Java Developer Jobs in NYC, NY
Java Full-Stack Engineer Jobs in NYC, NY
Javascript Jobs in NYC, NY
Lead Software Engineer Jobs in NYC, NY
Linux Jobs in NYC, NY
Manufacturing Engineer Jobs in NYC, NY
Mechanical Design Engineer Jobs in NYC, NY
Mechanical Engineering Jobs in NYC, NY
Mechatronics Engineering Jobs in NYC, NY
Network Engineer Jobs in NYC, NY
PHP Developer Jobs in NYC, NY
Platform Engineer Jobs in NYC, NY
Principal Architect Jobs in NYC, NY
Principal Engineer Jobs in NYC, NY
Principal Software Engineer Jobs in NYC, NY
Process Engineer Jobs in NYC, NY
Project Engineer Jobs in NYC, NY
Python Jobs in NYC, NY
QA Engineer Jobs in NYC, NY
QA Jobs in NYC, NY
Quality Assurance Automation Engineer Jobs in NYC, NY
Reliability Engineer Jobs in NYC, NY
Robotics Engineer Jobs in NYC, NY
Ruby Jobs in NYC, NY
Salesforce Developer Jobs in NYC, NY
Scala Jobs in NYC, NY
Security Engineer Jobs in NYC, NY
Software Engineer Jobs in NYC, NY
Solutions Architect Jobs in NYC, NY
Solutions Engineer Jobs in NYC, NY
SRE Engineer Jobs in NYC, NY
Staff Engineer Jobs in NYC, NY
Staff Software Engineer Jobs in NYC, NY
Systems Engineer Jobs in NYC, NY
Vice President Of Engineering Jobs in NYC, NY
Web Developer Jobs in NYC, NY
All Filters
Total selected ()
No Results
No Results




























