Top Senior Site Reliability Engineer Jobs in NYC, NY

Reposted 2 Days AgoSaved
In-Office
New York, NY
237K-321K Annually
Senior level
237K-321K Annually
Senior level
Artificial Intelligence • Legal Tech • Software
As a Senior Site Reliability Engineer, you'll operate foundational platform services, enhance reliability standards, automate processes, and work with engineering teams to improve systems.
Top Skills: Cloud InfrastructureKubernetesObservability Tools
Reposted 2 Days AgoSaved
Hybrid
New York, NY
140K-215K Annually
Senior level
140K-215K Annually
Senior level
Cloud • Computer Vision • Information Technology • Sales • Security • Cybersecurity
Lead architecture and implementation of reliability improvements across CrowdStrike's cloud-native platform. Build shared libraries and services, drive observability and SLO practices, perform performance and cost optimization, run resilience engineering and chaos experiments, automate infrastructure-as-code, mentor engineers, and embed with product teams to deliver scalable, highly reliable distributed systems at organizational scale.
Top Skills: AIAlertingAWSCassandraElasticsearchGCPGoInfrastructure-As-CodeJavaKafkaKotlinKubernetesNode.jsObservability (TracingOciOpensearchProfilingProtobufPythonScalaSlos)
Reposted 4 Days AgoSaved
Easy Apply
Remote or Hybrid
New York, NY
Easy Apply
127K-249K Annually
Senior level
127K-249K Annually
Senior level
Big Data • Cloud • Software • Database
Develop and maintain Kubernetes runtime environments, support developers, resolve critical issues, and participate in on-call rotations for production systems.
Top Skills: AWSAzureCert-ManagerCorednsCrdsCriCsiGatekeeperGCPGoHelmKubernetesKustomizeOperatorsPythonTerraform
6 Hours AgoSaved
Hybrid
New York, NY
200K-250K Annually
Expert/Leader
200K-250K Annually
Expert/Leader
Cloud • Healthtech • Internet of Things • Machine Learning • Software
Lead reliability, performance, automation, and scalability for a 24/7 AWS-based healthcare platform. Responsibilities include designing fault-tolerant infrastructure, managing Kubernetes and Docker environments, implementing Terraform and GitOps, improving observability, leading incident response and disaster recovery, and ensuring HIPAA and SOC 2 compliance. The role also sets technical strategy, reduces downtime through self-healing automation, and leads and mentors the SRE team.
Top Skills: AWSCi/CdDatadogDockerFhirGitopsGrafanaHipaaHl7KubernetesOpentelemetryPrometheusSoc 2Terraform
Reposted 9 Hours AgoSaved
In-Office
New York, NY
150K-190K Annually
Senior level
150K-190K Annually
Senior level
Fintech • Financial Services
Lead and improve production reliability for Wealth Management services: monitor availability and performance, automate deployments and operational tasks, troubleshoot infrastructure and databases, collaborate with development and upstream teams, maintain runbooks/knowledgebase, perform root cause analysis, and participate in 24/7 on-call coverage and offshore coordination to reduce outages and operational toil.
Top Skills: AutosysAWSAzureC#Ci/CdContainersDb2Ip SoftJavaJenkinsLinuxMqOraclePerlPythonRubyShellSockeyeSplunkSybaseTrainUnixVirtual MachinesWeb ServicesWindeployWindows
Reposted 11 Hours AgoSaved
Hybrid
New York, NY
95K-125K Annually
Mid level
95K-125K Annually
Mid level
Artificial Intelligence • eCommerce • Retail • Software
Build and maintain CI/CD pipelines, manage and automate cloud infrastructure and configurations, implement monitoring/logging and alerting for reliability, enforce security and compliance practices, and collaborate with development teams to support scaling and operations.
Top Skills: Soc ISoc Ii
Reposted 21 Hours AgoSaved
In-Office
New York, NY
160K-230K Annually
Mid level
160K-230K Annually
Mid level
Fintech • Payments • Financial Services
Operate and maintain AWS infrastructure and CI/CD pipelines, deploy containerized applications, write Terraform IaC, implement monitoring/observability, troubleshoot incidents, resolve security vulnerabilities, participate in on-call rotation, and produce operational/runbook documentation to ensure resilient cloud platform delivery.
Top Skills: AgileAlbAmazon Web Services (Aws)Aws IamAws LambdaCloudbeaverCloudwatchDevsecopsDirect ConnectDnsEc2EcsEfsEksElbGitlab CiGlueGrafanaHelmJavaJbossJIRANginxOktaOracle RdsPythonRds PostgresRoute53S3ScalaSnsTerraformTomcatTransit GatewayVpc
Reposted 19 Days AgoSaved
Remote
New York, NY
150K-200K Annually
Senior level
150K-200K Annually
Senior level
Artificial Intelligence • Cloud • Software • Infrastructure as a Service (IaaS)
Ensure stability and resilience of Runpod's distributed AI platform by defining SLIs/SLOs, leading incident response, building observability and reliability tooling, automating operational workflows, and partnering with engineering teams to reduce toil and improve production readiness.
Top Skills: BashCi/CdContainerized Production SystemsGoGpu Observability ToolingGrafanaInfrastructure As CodeLinuxPrometheusPython
Reposted 6 Days AgoSaved
Hybrid
New York, NY
147K-278K Annually
Senior level
147K-278K Annually
Senior level
Cloud • Software
Responsible for maintaining FedRAMP-compliant infrastructure, collaborating with software engineers, and ensuring system availability and security. Duties include infrastructure design, automation, monitoring, and incident response.
Top Skills: AWSGoKubernetesPuppetPythonTerraform
Reposted 6 Days AgoSaved
In-Office or Remote
New York, NY
161K-284K Annually
Senior level
161K-284K Annually
Senior level
Blockchain • eCommerce • Fintech • Payments • Software • Financial Services • Cryptocurrency
As a Senior Site Reliability Engineer, you will enhance platform reliability, lead incident management, and drive AI-driven improvements in operational workflows.
Top Skills: Amazon Web ServicesDatadogDynamoDBEnvoyEvent Driven ArchitecturesGrpcHTTPIstioJSONKotlinKubernetesLaunchdarklyModern JavaMySQLProtocol BuffersTerraformVitess
Reposted 20 Days AgoSaved
Easy Apply
Remote or Hybrid
New York, NY
Easy Apply
Internship
Internship
Cloud • Information Technology • Security • Software • Cybersecurity
This internship role focuses on SRE skills, requiring collaboration and problem-solving in dynamic environments for Zscaler's Zero Trust Exchange team.
Top Skills: AnsibleAws EcsKubernetesLinuxPythonTerraform
2 Days AgoSaved
In-Office
New York, NY
170K-225K Annually
Senior level
170K-225K Annually
Senior level
Fintech • Software • Financial Services
Own AWS and Kubernetes reliability, deployments, observability, incident response, infrastructure as code, CI/CD, and compliance controls for SOC 2 and PCI-DSS. Build scalable platform foundations, AI infrastructure, and greenfield product infrastructure while improving cost efficiency, security, vendor flexibility, and deployment consistency. Participate in on-call support and partner with security and engineering teams across backend, mobile, data, ML, and AI.
Top Skills: Ai/Llm ToolingAmazon EksAWSAws CdkCi/CdDatadogKotlinKubernetesMcpNode.jsOpentelemetryOpentofuPci-DssPythonSoc 2TerraformTypescript
New

Cut your apply time in half.

Use ourAI Assistantto automatically fill your job applications.

Use For Free
Application Tracker Preview
Reposted 2 Days AgoSaved
In-Office or Remote
New York, NY
110K-130K Annually
Mid level
110K-130K Annually
Mid level
Fintech • Software
Owner of operational health and availability for a FedRAMP High cloud platform. Act as L3 escalation, monitor observability, lead incident response, perform root cause analysis, automate toil, support deployments, maintain runbooks, and ensure compliance while partnering with engineering to improve reliability and resilience.
Top Skills: AWSAws BackupBashCi/CdCloudwatchDatadogEksGoGrafanaIamIstioJira Service ManagementKubernetesLinuxOpentelemetryPagerdutyPowershellPrometheusPythonRdsRoute 53SplunkVpc
Reposted 2 Days AgoSaved
In-Office or Remote
New York, NY
110K-120K Annually
Mid level
110K-120K Annually
Mid level
Fintech • Software
Operate and improve the reliability, availability, and security of a FedRAMP High cloud platform as an SRE/L3 escalation point. Monitor systems, troubleshoot incidents, lead incident response and RCA, build runbooks, enhance observability and automation, support deployments, and ensure compliance while partnering with engineering teams to reduce toil and improve resilience.
Top Skills: AWSAws BackupBashCi/CdCloudwatchConfiguration ManagementDatadogEksFedramp HighGoGrafanaIamInfrastructure As CodeIstioJira Service ManagementKubernetesLinuxOpentelemetryPagerdutyPowershellPrometheusPythonRdsRoute 53SplunkVpc
Reposted 2 Days AgoSaved
In-Office
New York, NY
220K-260K Annually
Expert/Leader
220K-260K Annually
Expert/Leader
Payments • Software • Automation
Lead platform and infrastructure direction on AWS, evolve CI/CD and ephemeral environments, set observability and SLO standards, drive incident response and postmortems, mentor engineers, and build automation to reduce operational risk.
Top Skills: AWSCi/CdDistributed SystemsEcsEphemeral Environments/Preview DeploysFargateGithub ActionsLogsObservability (MetricsSlos/Slis/Error BudgetsTracing)
Reposted 2 Days AgoSaved
In-Office or Remote
New York, NY
165K-215K Annually
Senior level
165K-215K Annually
Senior level
Software • Cybersecurity
This role involves managing Kubernetes clusters, cloud infrastructure, and CI/CD pipelines. The engineer will enhance system reliability and efficiency while troubleshooting production issues.
Top Skills: AlertmanagerAWSAzureBashCi/CdDockerElastic StackElasticsearchGCPGoGrafanaHelmKafkaKubernetesLokiMongoDBOciPrometheusPythonRedisSparkTerraform
Reposted 2 Days AgoSaved
In-Office
New York, NY
115K-125K Annually
Mid level
115K-125K Annually
Mid level
Fintech • Payments • Financial Services
The Site Reliability Engineer will assist clients with Redline products, manage production environments, troubleshoot issues, and ensure automation and customer satisfaction.
Top Skills: C/C++JavaLinuxPython
Reposted 2 Days AgoSaved
Remote or Hybrid
New York, NY
160K-180K Annually
Senior level
160K-180K Annually
Senior level
Artificial Intelligence • Machine Learning • Software • Analytics
The role involves end-to-end ownership of AWS infrastructure, managing Kubernetes platforms, and ensuring system reliability through observability and automation. Responsibilities include incident response and maintaining CI/CD systems.
Top Skills: ArgocdAWSDatadogGitGoKubernetesPythonTerraform
3 Days AgoSaved
Remote or Hybrid
New York, NY
230K-260K Annually
Senior level
230K-260K Annually
Senior level
Artificial Intelligence • Healthtech • Machine Learning • Software
Define reliability strategy and technical roadmaps; design scalable cloud infrastructure; improve observability, incident response, disaster recovery, automation, and deployment systems. Partner across engineering, security, data, AI, and product teams to strengthen resilience, compliance, and operational excellence. Lead architecture reviews, resolve complex production issues, establish SRE practices, and mentor engineers while reducing operational toil and improving platform scalability.
Top Skills: AutomationCi/CdCloud InfrastructureContainer OrchestrationContainersDisaster RecoveryDistributed SystemsInfrastructure As CodeMulti-Region ArchitectureNetworkingObservability
4 Days AgoSaved
In-Office
New York, NY
150K-160K Annually
Senior level
150K-160K Annually
Senior level
Natural Language Processing • Software • Conversational AI
Design, deploy, and manage resilient GCP network architectures, including VPCs, Interconnects, load balancing, and DNS. Lead Terraform-based infrastructure automation, implement network security policies, optimize cloud networking costs, and resolve connectivity and performance issues. Drive SRE practices such as capacity planning, error budgets, monitoring, and alerting. Partner with SRE, platform, and engineering teams to onboard services and maintain high availability, reliability, and security across distributed systems.
Top Skills: BashCloud InterconnectDnsGoGoogle Cloud Platform (Gcp)Infrastructure As Code (Iac)LinuxLoad BalancingPythonTerraformVpc
Reposted YesterdaySaved
Remote
New York, NY
180K-220K Annually
Senior level
180K-220K Annually
Senior level
Software • Defense
Work as an SRE embedded with product teams to improve reliability by fixing application code (primarily TypeScript), building observability (Prometheus, Loki, Grafana, Alloy), defining SLIs/SLOs, leading incident response and postmortems, automating toil, and supporting deployments across on‑prem DoD and AWS environments.
Top Skills: AlloyAWSBashContainersDockerGithub ActionsGitlab Ci/CdGoGrafanaJenkinsKubectlKubernetesLokiNode.jsPrometheusPythonTypescript
Reposted 13 Minutes AgoSaved
Remote
New York, NY
Internship
Internship
Artificial Intelligence • Cloud • Information Technology • Consulting
Internship SRE role responsible for availability, performance, and scalability of an e-commerce supply-chain platform. Tasks include SLO/SLA definition, observability (Prometheus/Grafana/Loki/Tempo/OpenTelemetry), incident response, capacity planning, disaster recovery for PostgreSQL, infrastructure-as-code (Terraform), CI/CD automation, and operational reliability for AI agent services. Mentored by Head of Technology/CTO with potential conversion to full-time based on performance.
Top Skills: BashCi/CdDockerGrafanaLangchainLlmLokiMakefileNestjsOpentelemetryOracle CloudPgbackrestPostgresql 15PrometheusPythonRedisTempoTerraformTraefik
Reposted 9 Days AgoSaved
In-Office or Remote
New York, NY
135K-160K Annually
Senior level
135K-160K Annually
Senior level
Artificial Intelligence • Healthtech • Software • Telehealth
Design, deploy, and maintain AWS-hosted Kubernetes (EKS) infrastructure; build automation and AI-assisted runbooks; provide observability and incident response; ensure HIPAA-compliant, high-availability platform operations and mentor engineering teams.
Top Skills: AWSBashDatadogEc2EksGitGithub ActionsGoHelmKubernetesPythonRdsS3Terraform
Reposted 9 Days AgoSaved
In-Office
New York, NY
208K-269K Annually
Senior level
208K-269K Annually
Senior level
Artificial Intelligence • Software
Own and scale the GPU compute fleet: build metrics, alerting, observability, and repair automation; design GPU qualification/burn-in pipelines; own Redfish/BMC tooling and firmware telemetry; run incidents, eliminate toil, and deliver end-to-end reliability and orchestration for Kubernetes and bare-metal compute at hyperscale.
Top Skills: Agentic FrameworksBare MetalBmcCadenceClaude CodeCursorFirmwareGoGpuGrafanaIpmiKubernetesLlm ApisMcp ServersPrometheusPythonRedfishTemporal
Reposted 21 Hours AgoSaved
Remote
New York, NY
115K-135K Annually
Mid level
115K-135K Annually
Mid level
Aerospace • Manufacturing
As a Site Reliability Engineer, you'll build and manage observability platforms for satellite communications, define SLOs/SLIs, and collaborate on incident response and deployment automation.
Top Skills: ArgocdAWSElkGCPGoGrafanaIstioJaegerKubernetesLinkerdLokiOpentelemetryPrometheusPythonTempoTerraform
All Filters
JobType
New Jobs
Job Category
Experience
Industry
Company Name
Company Size

Sign up now Access later

Create Free Account