Top Reliability Engineer Jobs in NYC, NY

YesterdaySaved
In-Office
New York, NY
150K-225K Annually
Senior level
150K-225K Annually
Senior level
Consulting • Quantitative Trading
Build and operate a reliable, scalable internal Agentic AI platform. Define SLOs, error budgets, observability standards, incident response runbooks, and reliability practices. Develop automation and Python applications, integrate APIs and asynchronous systems, manage cloud and Kubernetes deployments, and support users during incidents. Work with databases, CI/CD pipelines, AI agents, and RAG workflows while improving code quality, developer processes, platform performance, availability, accuracy, and cost efficiency.
Top Skills: Ai AgentsAmazon OpensearchAmazon S3Asynchronous ArchitectureAWSCi/CdDatadogDynamoDBElasticsearchEvent-Driven ArchitectureGithub ActionsKubernetesMcp ServersMySQLPostgresPythonRest ApisRetrieval-Augmented Generation (Rag)
Reposted 23 Days AgoSaved
In-Office
New York, NY
160K-220K Annually
Senior level
160K-220K Annually
Senior level
Cloud
The role involves designing, optimizing, and maintaining PostgreSQL and MySQL databases, ensuring high availability, reliability, and performance for mission-critical systems, while automating operational tasks and responding to incidents.
Top Skills: AnsibleAWSDatadogGCPGoGrafanaKubernetesMySQLPostgresPrometheusPythonTerraform
Reposted YesterdaySaved
Hybrid
New York, NY
95K-125K Annually
Mid level
95K-125K Annually
Mid level
Artificial Intelligence • eCommerce • Retail • Software
Build and maintain CI/CD pipelines, manage and automate cloud infrastructure and configurations, implement monitoring/logging and alerting for reliability, enforce security and compliance practices, and collaborate with development teams to support scaling and operations.
Top Skills: Soc ISoc Ii
Reposted 24 Days AgoSaved
In-Office
New York, NY
160K-230K Annually
Mid level
160K-230K Annually
Mid level
Fintech • Payments • Financial Services
Operate and maintain AWS infrastructure and CI/CD pipelines, deploy containerized applications, write Terraform IaC, implement monitoring/observability, troubleshoot incidents, resolve security vulnerabilities, participate in on-call rotation, and produce operational/runbook documentation to ensure resilient cloud platform delivery.
Top Skills: AgileAlbAmazon Web Services (Aws)Aws IamAws LambdaCloudbeaverCloudwatchDevsecopsDirect ConnectDnsEc2EcsEfsEksElbGitlab CiGlueGrafanaHelmJavaJbossJIRANginxOktaOracle RdsPythonRds PostgresRoute53S3ScalaSnsTerraformTomcatTransit GatewayVpc
Reposted One Month AgoSaved
Hybrid
New York, NY
149K-282K Annually
Senior level
149K-282K Annually
Senior level
Cloud • Software
Responsible for maintaining FedRAMP-compliant infrastructure, collaborating with software engineers, and ensuring system availability and security. Duties include infrastructure design, automation, monitoring, and incident response.
Top Skills: AWSGoKubernetesPuppetPythonTerraform
Reposted One Month AgoSaved
In-Office or Remote
New York, NY
161K-284K Annually
Senior level
161K-284K Annually
Senior level
Blockchain • eCommerce • Fintech • Payments • Software • Financial Services • Cryptocurrency
As a Senior Site Reliability Engineer, you will enhance platform reliability, lead incident management, and drive AI-driven improvements in operational workflows.
Top Skills: Amazon Web ServicesDatadogDynamoDBEnvoyEvent Driven ArchitecturesGrpcHTTPIstioJSONKotlinKubernetesLaunchdarklyModern JavaMySQLProtocol BuffersTerraformVitess
Reposted One Month AgoSaved
In-Office or Remote
New York, NY
105K-300K Annually
Entry level
105K-300K Annually
Entry level
Information Technology • Software • Financial Services • Big Data Analytics
SREs at Citadel focus on optimizing and maintaining system reliability, performance, and automation for investment applications, collaborating closely with teams.
Top Skills: Ci/CdCSSJavaScriptPythonReactSQL
3 Days AgoSaved
In-Office
New York, NY
120K-180K Annually
Mid level
120K-180K Annually
Mid level
Professional Services • Consulting
Own end-to-end infrastructure as the company’s first dedicated SRE. Design, build, scale, and operate cloud, on-premises, hybrid, and air-gapped environments; architect AWS migrations; support customer-specific deployments; and improve observability, networking, containerization, CI/CD, and delivery systems. The role is full-time and on-site in Midtown Manhattan, with relocation support but no visa sponsorship.
Top Skills: AWSAzureCi/CdDockerGCPGitopsKubernetesNetworkingObservability ToolingOpentofuTerraformTerragrunt
Reposted One Month AgoSaved
Hybrid
New York, NY
Mid level
Mid level
Financial Services
Operate and improve production reliability for critical services by building automation, monitoring, and runbooks. Triage incidents, reduce MTTR, improve observability, partner with engineering for root-cause fixes, and apply validated enterprise AI-assisted tools to support SRE workflows and reduce toil.
Top Skills: .NetAWSCi/CdDatadogEnterprise-Authorized AiJavaKafkaMqPythonSpring Boot
27 Days AgoSaved
Remote or Hybrid
New York, NY
Senior level
Senior level
AdTech • Software
Leads strategy, architecture, deployment, operations, automation, observability, and reliability initiatives for globally distributed database systems. Partners across engineering and business teams to improve performance, scalability, security, continuity, and operational rigor. Oversees incident troubleshooting, root-cause analyses, release processes, database change management, technology evaluation, roadmap planning, and mentoring while operating databases across Kubernetes, cloud, and bare-metal environments.
Top Skills: AerospikeAlembicAnsibleArgocdBashCassandraCassandraEtcdFlinkFlywayGaleraGitlabGoGrafanaHadoopHelmIcebergJavaKafkaKotlinKubernetesLiquibaseLokiMariadbMySQLPostgresPrometheusPythonRedisScylladbShellSparkStarrocksTerraformTrinoVerticaZookeeper
4 Days AgoSaved
In-Office
New York, NY
175K-200K Annually
Mid level
175K-200K Annually
Mid level
Software
Site Reliability Engineer responsible for improving platform reliability, observability, and security. Duties include expanding PagerDuty on-call operations, implementing Datadog metrics and tracing, and managing credentials, certificates, encryption keys, and their rotation. The role supports infrastructure handling protected health information and offers mentorship toward broader SRE ownership.
Top Skills: Aws AcmAws KmsAws Secrets ManagerDatadogHashicorp VaultPagerduty
Reposted One Month AgoSaved
Hybrid
New York, NY
112K-159K Annually
Senior level
112K-159K Annually
Senior level
Artificial Intelligence • Cloud • Fintech • Information Technology • Insurance • Financial Services • Big Data Analytics
Lead application readiness and remediation coordination for AWS, EOL, and vulnerability patches. Validate impacts, define smoke and regression tests, drive automation, resolve dependencies, escalate blockers, and secure production sign-off to ensure audit-ready closure.
Top Skills: AmiApi TestingAWSCertificatesCi/Cd PipelinesContainerizationDastDatabasesDockerEc2EksLibrariesMiddlewareNew RelicNew Relic MonitorsObservability ToolingRegression AutomationRuntimesSastScaService DashboardsSmoke TestingSyntheticsTerraform
New

Track Smarter, Apply Better.

Ditch the spreadsheets. Organize your job search with our freeApplication Tracker.

Use For Free
Application Tracker Preview
6 Days AgoSaved
In-Office
New York, NY
120K-150K Annually
Entry level
120K-150K Annually
Entry level
Financial Services
Supports reliability and operations for hybrid on-premises and multi-cloud infrastructure across Azure, AWS, and GCP. Monitors production systems, responds to incidents, maintains compute, storage, networking, middleware, and identity services, and contributes to Infrastructure-as-Code, DevSecOps, GitOps, CI/CD, cloud security, observability, documentation, and post-incident reviews. Participates in on-call rotations and helps improve platform performance, resilience, and consistency.
Top Skills: AnsibleAWSAzureBashCi/CdDatadogDevsecopsDhcpDnsDockerFirewallsGCPGitopsGrafanaKubernetesLinuxPowershellPrometheusPythonTcp/IpTerraformVlansWindows Server
Reposted One Month AgoSaved
Easy Apply
Hybrid
New York, NY
Easy Apply
151K-191K Annually
Senior level
151K-191K Annually
Senior level
Fintech • Information Technology • Software • Financial Services
The role involves designing and automating infrastructure management, improving reliability, building internal tools, and contributing to architectural decisions. Responsibilities include working with Kubernetes and managing large-scale infrastructure, while participating in on-call rotations to prevent incidents.
Top Skills: CloudwatchDatadogDockerEfkElkGoJavaScriptKubernetesPythonTerraform
Reposted One Month AgoSaved
In-Office or Remote
New York, NY
150K-250K Annually
Mid level
150K-250K Annually
Mid level
Mobile • Software
Site Reliability Engineers will work on production infrastructure, focusing on AWS and Kubernetes while ensuring high availability and customer satisfaction.
Top Skills: AirflowAWSCircleCICloudwatchEksGrafanaMongoDBPagerdutyPingdomRustScala SparkTerraformTypescript
Reposted One Month AgoSaved
Easy Apply
Remote or Hybrid
New York, NY
Easy Apply
126K-248K Annually
Senior level
126K-248K Annually
Senior level
Big Data • Cloud • Software • Database
The Senior Site Reliability Engineer will develop and support distributed storage services, ensuring reliability and operational safety, with a focus on automation and efficiency.
Top Skills: AWSAzureDnsGoGoogle Cloud PlatformKubernetesLinuxPythonTcp/IpTls
Reposted One Month AgoSaved
Easy Apply
Remote or Hybrid
New York, NY
Easy Apply
127K-249K Annually
Expert/Leader
127K-249K Annually
Expert/Leader
Big Data • Cloud • Software • Database
Seeking a Site Reliability Engineer with expertise in networking and distributed systems for building secure multi-cloud infrastructure. Responsibilities include maintaining network architecture and ensuring reliable service-to-service communication, involving a 24/7 on-call rotation.
Top Skills: AWSAzureBgpDnsGCPIpv6KubernetesLoad BalancingMtlsService MeshTcp/IpTlsVpcsVpns
180K-220K Annually
Senior level
Software • Defense
Work as an SRE embedded with product teams to improve reliability by fixing application code (primarily TypeScript), building observability (Prometheus, Loki, Grafana, Alloy), defining SLIs/SLOs, leading incident response and postmortems, automating toil, and supporting deployments across on‑prem DoD and AWS environments.
Top Skills: AlloyAWSBashContainersDockerGithub ActionsGitlab Ci/CdGoGrafanaJenkinsKubectlKubernetesLokiNode.jsPrometheusPythonTypescript
Reposted 8 Days AgoSaved
Hybrid
New York, NY
180K-200K Annually
Senior level
180K-200K Annually
Senior level
Productivity
Build and operate reliable, scalable infrastructure and backend systems. Define and monitor SLOs and SLAs, manage on-call operations, improve observability, performance, security, developer experience, and technical processes. Contribute to Temporal-based workflow orchestration, advise backend projects, manage AWS infrastructure, and support system scalability as the customer base grows.
Top Skills: Amazon EksAWSCi/CdDockerHelmInfrastructure As CodeKubernetesLinuxTemporal
One Month AgoSaved
Remote or Hybrid
New York, NY
Senior level
Senior level
AdTech • Cloud • Digital Media • Information Technology • News + Entertainment • App development
Lead SRE for cloud-based live linear playout systems driving reliability, observability, incident response, SLIs/SLOs, automation, capacity planning, runbooks, and L1/L2 on-call support to ensure resilient distribution across NBCUniversal channels.
Top Skills: AmagiAWSCmafDockerEsamGrafanaH.264HarrisHevcHlsImagineIp NetworkingKubernetesLinuxMicrosoft TeamsRistScte-224Scte-35ServicenowSlackSnellSplunkSrtTs
22 Days AgoSaved
Remote
New York, NY
220K-240K Annually
Senior level
220K-240K Annually
Senior level
Software
Own the reliability, performance, availability, scalability, and cost efficiency of production Aurora MySQL databases on AWS. Responsibilities include observability, alerting, query optimization, replication, backups, disaster recovery, schema migrations, reader topology, cost optimization, runbooks, incident response, and database performance reviews. The role partners with engineering teams, supports HIPAA-compliant handling of PHI, and establishes reliable database operating practices.
Top Skills: Amazon RdsAmazon RedshiftAurora MysqlAWSBashDatabricksDatadogGrafanaMySQLPercona Monitoring And Management (Pmm)Percona ToolkitPerformance InsightsPHPPrometheusPythonSnowflakeTerraform
An Hour AgoSaved
Remote
New York, NY
101K-161K Annually
Senior level
101K-161K Annually
Senior level
Cloud • Software • Analytics
Develop and operate Arista’s FedRAMP CloudVision SaaS platform at scale. Responsibilities include improving reliability, scalability, observability, autoscaling, disaster recovery, capacity planning, CI/CD, network architecture, cost optimization, and cloud application security. The role develops and manages Kubernetes-native services, distributed databases, and automation using technologies such as GCP, GKE, Go, Python, Ansible, Pulumi, and Bash. Participation in a FedRAMP on-call rotation is required.
Top Skills: AnsibleBashCi/CdDistributed DatabasesFedrampGoGoogle Cloud Platform (Gcp)Google Kubernetes Engine (Gke)KubernetesPulumiPythonSaaS
Reposted One Month AgoSaved
In-Office
New York, NY
125K-350K Annually
Mid level
125K-350K Annually
Mid level
Information Technology • Software • Financial Services • Quantitative Trading
The Site Reliability Engineer will provide support and diagnose issues within a real-time, distributed environment, focusing on large-scale application and infrastructure management, with basic required skills in UNIX/Linux, networking, SQL, and scripting languages.
Top Skills: BashPythonSQLTcp/IpUdpUnix/Linux
Reposted 8 Hours AgoSaved
Remote
New York, NY
120K-165K Annually
Senior level
120K-165K Annually
Senior level
Fitness • Healthtech • Software
Own reliability and security of CI/CD and production services: define SLI/SLOs, lead incident response and postmortems, build observability (Datadog), operate Kubernetes and IaC (Terraform), harden pipelines with SAST/DAST/SCA and policy-as-code, and coach teams on reliability and operational best practices.
Top Skills: AWSCi/CdConftestDastDatadogDockerGithub ActionsGoInfrastructure As CodeKubernetesKyvernoOpa/RegoPythonSastScaTerraformTypescript
11 Hours AgoSaved
Remote
New York, NY
140K-170K Annually
Senior level
140K-170K Annually
Senior level
eCommerce • Manufacturing
Manage Azure infrastructure and AKS clusters, build GitHub Actions CI/CD pipelines, and improve Grafana-based observability and incident response. Define SLIs, SLOs, and error budgets; maintain infrastructure as code with Pulumi; troubleshoot reliability issues; perform capacity planning and performance tuning; participate in on-call support; and document operational procedures. Collaborate with development teams to deliver scalable, reliable production systems.
Top Skills: Azure Kubernetes Service (Aks)BashDnsGithub ActionsGrafanaIstioKubernetesLoad BalancingLokiAzurePrometheusPulumiPython
All Filters
JobType
New Jobs
Job Category
Experience
Industry
Company Name
Company Size

Sign up now Access later

Create Free Account