Maximum of 25 job preferences reached.
Top Senior Site Reliability Engineer Jobs in NYC, NY
Hardware • Machine Learning • Security • Software
Design and maintain developer experience tooling and CI/CD pipelines, manage self-service deployment tools (secrets, rollbacks), partner with cloud teams to troubleshoot production systems, participate in on-call rotation, and write production-grade automation and tests in Go and TypeScript to improve developer velocity and reliability.
Top Skills:
AWSGithub ActionsGoHelmKubernetesTerraformTypescript
Cybersecurity
Assist with monitoring system performance, uptime, and reliability; support incident response, troubleshooting, and root cause analysis; monitor SRE alerts in Slack and report issues; learn cloud platforms; and collaborate with development and DevOps teams.
Top Skills:
AWSAzureDevOpsDevsecopsGCPGrafanaKubernetesPrometheusSlack
Fintech • Software
Lead SRE efforts for DFIN SaaS: ensure availability, performance, scalability, and automation. Implement monitoring, CI/CD, IaC, container orchestration, AI-enhanced observability, incident response, RCA, and runbook automation while collaborating across engineering teams.
Top Skills:
.NetAiopsAksAnsibleAppdynamicsAWSAzureAzure DevopsBashC#Ci/CdCloud Ai ServicesContainersCosmosDatadogDynatraceEksFirewallHarnessIdera Sql Diagnostic ManagerInfrastructure As Code (Iac)JavaJenkinsKubernetesLinuxLoad BalancingNew RelicPowershellPythonRedgate Sql MonitorSolarwinds Database Performance AnalyzerSQLTerraformWindows
Information Technology • Consulting
As a Senior Staff Site Reliability Engineer, you will lead the SRE team, advocate best practices, ensure resilience in cloud architecture, and mentor team members.
Top Skills:
ArgocdCircleCIGoogle Cloud PlatformKubernetesPulumiTerraformTypescript
Cloud • Security • Software • Cybersecurity
Deploy and operate scalable, highly available cloud systems; improve application and network security, reliability, speed, and capacity; automate cloud deployments; monitor and troubleshoot services to meet SLAs; analyze logs and events; and resolve infrastructure issues across large-scale cloud infrastructure.
Top Skills:
AnsibleAWSBashBitbucketDatabricksDatadogDigicertGradleJenkinsNode.jsPythonSplunkTenableTerraformThreatmetrix
Artificial Intelligence • Fintech • Machine Learning • Natural Language Processing • Business Intelligence
Lead architecture and implementation of reliability platforms and SRE practices for a production SaaS. Build self-service reliability tooling, drive AIOps automation, advance observability (monitoring, tracing, profiling), lead incident response and postmortems, mentor engineers, and embed production readiness across teams to achieve 99.99% uptime.
Top Skills:
AWSAzureContinuous ProfilingDatadogDnsElkGCPGoGrafanaHttp/SKubernetesLoad BalancingOpentelemetryPrometheusPythonTcp/Ip
Artificial Intelligence • Fintech • Machine Learning • Software • App development • Conversational AI • Generative AI
Own and improve production infrastructure reliability, deployments, Infrastructure-as-Code, Kubernetes environments, automation, CI/CD, monitoring, alerting, and observability. Investigate incidents, optimize system performance, maintain documentation and runbooks, and support DNS, WAF, CDN, and caching infrastructure. The role requires strong Linux administration, Bash scripting, networking, Git, and containerization skills, with independent ownership and collaboration across development and operations teams.
Top Skills:
AkamaiAmqpAnsibleAWSBashCdnCloudflareDnsDockerGCPGitGitlab CiGrafanaHttp/HttpsKubernetesLinuxPodmanPrometheusPythonRabbitMQTerraformVictoriametricsWafZabbix
Automotive
Design and implement scalable cloud infrastructure, monitor performance, automate processes, ensure security and compliance, and lead a DevOps team.
Top Skills:
AWSBashCi/CdDockerElk StackGCPGrafanaKubernetesPrometheusPythonTerraform
Cloud • Software
The Site Reliability Engineer will ensure reliable cloud operations by applying Python for infrastructure automation, managing OpenStack and Kubernetes, and practicing devsecops in a fast-paced environment.
Top Skills:
KubernetesLinuxOpenstackPython
Cloud • Security • Software • Cybersecurity
As a Site Reliability Engineer II, you'll automate tasks, monitor AI workloads, enhance dashboards, support CI/CD processes, and collaborate with engineering teams on complex issues while participating in on-call rotations.
Top Skills:
GoGrafanaKubernetesLinuxPrometheusPythonSaltstackTerraform
Software • Web3
Lead reliability practices across teams: embed early in projects, define SLIs/SLOs, build multi-cloud paved roads with Terraform, run on-call, drive org-wide incident maturity and tooling.
Top Skills:
AWSAzureGCPRuby On RailsTerraformTypescriptWebcontainers
Blockchain • Software
Build, operate, and scale production Kubernetes infrastructure using GitOps and declarative IaC. Design CI/CD workflows, observability, and secure-by-default systems. Troubleshoot networking/storage, participate in on-call rotations, automate operational workflows, and drive postmortems and reliability improvements.
Top Skills:
ArbitrumArgocdArgocd ApplicationsetsAWSAzureBashCloudwatchCodebuildGCPGithub ActionsGitopsGoGrafanaK9SKubernetesLinuxLokiMimirPrometheusPrysmPythonTerraformYamlZerodev
New
Cut your apply time in half.
Use ourAI Assistantto automatically fill your job applications.
Use For Free
Artificial Intelligence
Build and run AI-native GTM systems: design and execute outbound experiments, build ICP scoring and intent pipelines, maintain prospecting/enrichment stack and CRM automations, extend LLM/agent workflows, and deliver dashboards tying activity to ARR.
Top Skills:
Agent FrameworksAPIsApolloClayCRMHubspotInstantlyJavaScriptLemlistLinkedInLlm ApisPythonSQL
Software
As an AI Support Engineer, you'll manage support requests, resolve user issues, optimize ML models, and contribute to product development.
Top Skills:
Tensorrt
Fintech
Lead adoption of SRE practices to improve reliability, observability, automation, and incident response. Implement and maintain observability tooling, instrumentation, CI/CD, and infrastructure-as-code. Partner with developers, participate in on-call rotations, drive postmortems, and reduce operational overhead through automation.
Top Skills:
AnthropicAWSAws EcsAws EksAzureC#DockerGitlab CiGrafanaLinuxOpenaiPrometheusPuppetPythonSplunkTerraformTypescriptWindows
Fitness • Healthtech • Software
Own reliability and security of CI/CD and production services: define SLI/SLOs, lead incident response and postmortems, build observability (Datadog), operate Kubernetes and IaC (Terraform), harden pipelines with SAST/DAST/SCA and policy-as-code, and coach teams on reliability and operational best practices.
Top Skills:
AWSCi/CdConftestDastDatadogDockerGithub ActionsGoInfrastructure As CodeKubernetesKyvernoOpa/RegoPythonSastScaTerraformTypescript
Automotive
Leads SRE engineering leaders and engineers while defining enterprise observability, reliability, and platform strategy across GCP, on-premise, manufacturing, distribution, and campus environments. Oversees vendor-agnostic tooling, OpenTelemetry integrations, CI/CD observability, SRE maturity models, and Agentic AI initiatives. Drives adoption of SRE practices, develops technical roadmaps, partners with senior leadership and operational teams, and maintains hands-on architectural and technical credibility.
Top Skills:
Agentic AiAWSAzureCi/CdDatadogDynatraceGCPNew RelicOpentelemetryOtel Genai Semantic ConventionsSource Control PlatformsSplunkTerraform
Artificial Intelligence • Natural Language Processing • Generative AI
Lead deployment and runtime operations for ML safety systems: configure and verify safeguards across platforms, run canary rollouts and validations, automate launch runbooks into pipelines, maintain a provenance-backed safeguards registry, and participate in on-call and incident response for safety-critical model releases.
Top Skills:
AWSAws BedrockCanary DeploymentsCi/Cd PipelinesClaudeConfig ManagementGCPGcp VertexLlm InferencePythonRustTransformer-Based Models
Healthtech • Software
Own reliability, security, and continuity of Medgen production environments. Manage Windows and Linux servers, HA configurations, backups, DR, deployments, and CI/CD automation. Coordinate vulnerability assessments, incident response, cloud/hybrid evaluations, and infrastructure modernization.
Top Skills:
.NetAWSAzureBackupsCi/CdDisaster RecoveryDnsFirewallsGCPIisJavaScriptLinuxMicrosoft TfsNetwork SegmentationSQLTcp/IpVpnsWindows Server
Healthtech • Software
Own reliability, security, and continuity of Labgen production: manage Windows/Linux servers, HA, networking, backups/DR, production deployments, CI/CD, monitoring, cloud evaluation, incident response, vendor relations, and compliance.
Top Skills:
.NetApacheAWSAzureBashCentosCertificates/PkiCvsDatadogDnsElkFirewallsGCPGitGrafanaJavaScriptJenkinsLinux (RhelPrometheusPythonSentineloneSql/Relational DatabasesTcp/IpUbuntu)VpnWindows Server
Other
Design, build, and maintain highly available cloud-native systems. Improve reliability through automation, CI/CD, Kubernetes, observability, and incident management. Collaborate with developers, security, and product teams to define SLOs, implement self-healing, debug production issues, and ensure secure deployments.
Top Skills:
AWSAzure Cloud ServicesDatadogGCPGithub ActionsGitlab CiGoInfrastructure As CodeKubernetesOpsgeniePagerdutyPythonRubySite Reliability Engineering Foundation
Cloud • Security • Software • Cybersecurity
Design, develop, test, and operate scalable infrastructure and services for Akamai Cloud. Implement and manage Infrastructure-as-Code (Terraform and similar tools), CI/CD, and observability. Automate reliability improvements, mentor engineers, collaborate on incident response and root-cause remediation, and participate in on-call rotations.
Top Skills:
Alerting)AnsibleChefCi/CdInfrastructure As CodeLinuxLoggingObservability (MonitoringPuppetSaltstackTerraform
Artificial Intelligence • Information Technology • Consulting
Build and operate Nebius's network infrastructure: define SLIs/SLOs, improve site and inter-site reliability, lead incident response and postmortems, develop observability and alerting, automate change workflows, and collaborate with network and platform teams to embed operability.
Top Skills:
Ci/CdContainer PlatformsGoInfrastructure As CodeLinuxPython
Information Technology • Consulting
Designs and maintains AWS cloud infrastructure, CI/CD pipelines, Docker workloads, and IaC provisioning. Implements monitoring, logging, alerting, observability, security hardening, and code quality checks. Troubleshoots production issues, supports incident response and root-cause analysis, improves system reliability, and automates operational tasks. Collaborates with distributed engineering teams to establish infrastructure, deployment, monitoring, and SRE best practices.
Top Skills:
AWSAws IamCi/CdCloudFormationCloudwatchDatadogDockerGrafanaInfrastructure As Code (Iac)KubernetesPrometheusSonarqubeTerraform
Fintech • Financial Services
The Director of Splunk Platform Engineering & SRE owns the enterprise Splunk platform, drives incident resolution, optimizes systems, and mentors engineers, focusing on automation and performance.
Top Skills:
AnsibleGitGoJavaKubernetesLinux/UnixMoogPrometheusPythonSplunk
Let Your Resume Do The Work
Upload your resume to be matched with jobs you're a great fit for.
Success! We'll use this to further personalize your experience.
Top NYC Companies Hiring Senior Site Reliability Engineers
See AllPopular Job Searches
All Software Engineer Jobs in NYC
.NET Developer Jobs in NYC
Android Developer Jobs in NYC
C# Jobs in NYC
C++ Jobs in NYC
DevOps Jobs in NYC
Engineering Manager Jobs in NYC
Front End Developer Jobs in NYC
Golang Jobs in NYC
Hardware Engineer Jobs in NYC
iOS Developer Jobs in NYC
Java Developer Jobs in NYC
Javascript Jobs in NYC
Linux Jobs in NYC
Perl Jobs in NYC
PHP Developer Jobs in NYC
Python Jobs in NYC
QA Jobs in NYC
Ruby Jobs in NYC
Sales Engineer Jobs in NYC
Salesforce Developer Jobs in NYC
Scala Jobs in NYC
Artificial Intelligence Jobs in NYC
Artificial Intelligence Engineer Jobs in NYC
AWS Engineer Jobs in NYC
Backend Engineer Jobs in NYC
DevOps Engineer Jobs in NYC
Director of Engineering Jobs in NYC
Engineering Jobs in NYC
Full Stack Engineer Jobs in NYC
Infrastructure Engineer Jobs in NYC
Lead Software Engineer Jobs in NYC
Network Engineer Jobs in NYC
Platform Engineer Jobs in NYC
Principal Architect Jobs in NYC
Principal Engineer Jobs in NYC
Principal Software Engineer Jobs in NYC
Quality Assurance Automation Engineer Jobs in NYC
Reliability Engineer Jobs in NYC
Senior Backend Engineer Jobs in NYC
Senior Cloud Engineer Jobs in NYC
Senior Full-Stack Engineer Jobs in NYC
Senior Platform Engineer Jobs in NYC
Senior Python Engineer Jobs in NYC
Senior Site Reliability Engineer Jobs in NYC
Solutions Architect Jobs in NYC
Solutions Engineer Jobs in NYC
Staff Engineer Jobs in NYC
Staff Software Engineer Jobs in NYC
Systems Engineer Jobs in NYC
Vice President of Engineering Jobs in NYC
All Filters
Total selected ()
No Results
No Results
.png)





















.png)






