Maximum of 25 job preferences reached.
Top Reliability Engineer Jobs in NYC, NY
Consulting • Quantitative Trading
Build and operate a reliable, scalable internal Agentic AI platform. Define SLOs, error budgets, observability standards, incident response runbooks, and reliability practices. Develop automation and Python applications, integrate APIs and asynchronous systems, manage cloud and Kubernetes deployments, and support users during incidents. Work with databases, CI/CD pipelines, AI agents, and RAG workflows while improving code quality, developer processes, platform performance, availability, accuracy, and cost efficiency.
Top Skills:
Ai AgentsAmazon OpensearchAmazon S3Asynchronous ArchitectureAWSCi/CdDatadogDynamoDBElasticsearchEvent-Driven ArchitectureGithub ActionsKubernetesMcp ServersMySQLPostgresPythonRest ApisRetrieval-Augmented Generation (Rag)
Cloud
The role involves designing, optimizing, and maintaining PostgreSQL and MySQL databases, ensuring high availability, reliability, and performance for mission-critical systems, while automating operational tasks and responding to incidents.
Top Skills:
AnsibleAWSDatadogGCPGoGrafanaKubernetesMySQLPostgresPrometheusPythonTerraform
Artificial Intelligence • eCommerce • Retail • Software
Build and maintain CI/CD pipelines, manage and automate cloud infrastructure and configurations, implement monitoring/logging and alerting for reliability, enforce security and compliance practices, and collaborate with development teams to support scaling and operations.
Top Skills:
Soc ISoc Ii
Fintech • Payments • Financial Services
Operate and maintain AWS infrastructure and CI/CD pipelines, deploy containerized applications, write Terraform IaC, implement monitoring/observability, troubleshoot incidents, resolve security vulnerabilities, participate in on-call rotation, and produce operational/runbook documentation to ensure resilient cloud platform delivery.
Top Skills:
AgileAlbAmazon Web Services (Aws)Aws IamAws LambdaCloudbeaverCloudwatchDevsecopsDirect ConnectDnsEc2EcsEfsEksElbGitlab CiGlueGrafanaHelmJavaJbossJIRANginxOktaOracle RdsPythonRds PostgresRoute53S3ScalaSnsTerraformTomcatTransit GatewayVpc
Reposted One Month AgoSaved
Cloud • Software
Responsible for maintaining FedRAMP-compliant infrastructure, collaborating with software engineers, and ensuring system availability and security. Duties include infrastructure design, automation, monitoring, and incident response.
Top Skills:
AWSGoKubernetesPuppetPythonTerraform
Blockchain • eCommerce • Fintech • Payments • Software • Financial Services • Cryptocurrency
As a Senior Site Reliability Engineer, you will enhance platform reliability, lead incident management, and drive AI-driven improvements in operational workflows.
Top Skills:
Amazon Web ServicesDatadogDynamoDBEnvoyEvent Driven ArchitecturesGrpcHTTPIstioJSONKotlinKubernetesLaunchdarklyModern JavaMySQLProtocol BuffersTerraformVitess
Information Technology • Software • Financial Services • Big Data Analytics
SREs at Citadel focus on optimizing and maintaining system reliability, performance, and automation for investment applications, collaborating closely with teams.
Top Skills:
Ci/CdCSSJavaScriptPythonReactSQL
Professional Services • Consulting
Own end-to-end infrastructure as the company’s first dedicated SRE. Design, build, scale, and operate cloud, on-premises, hybrid, and air-gapped environments; architect AWS migrations; support customer-specific deployments; and improve observability, networking, containerization, CI/CD, and delivery systems. The role is full-time and on-site in Midtown Manhattan, with relocation support but no visa sponsorship.
Top Skills:
AWSAzureCi/CdDockerGCPGitopsKubernetesNetworkingObservability ToolingOpentofuTerraformTerragrunt
Financial Services
Operate and improve production reliability for critical services by building automation, monitoring, and runbooks. Triage incidents, reduce MTTR, improve observability, partner with engineering for root-cause fixes, and apply validated enterprise AI-assisted tools to support SRE workflows and reduce toil.
Top Skills:
.NetAWSCi/CdDatadogEnterprise-Authorized AiJavaKafkaMqPythonSpring Boot
AdTech • Software
Leads strategy, architecture, deployment, operations, automation, observability, and reliability initiatives for globally distributed database systems. Partners across engineering and business teams to improve performance, scalability, security, continuity, and operational rigor. Oversees incident troubleshooting, root-cause analyses, release processes, database change management, technology evaluation, roadmap planning, and mentoring while operating databases across Kubernetes, cloud, and bare-metal environments.
Top Skills:
AerospikeAlembicAnsibleArgocdBashCassandraCassandraEtcdFlinkFlywayGaleraGitlabGoGrafanaHadoopHelmIcebergJavaKafkaKotlinKubernetesLiquibaseLokiMariadbMySQLPostgresPrometheusPythonRedisScylladbShellSparkStarrocksTerraformTrinoVerticaZookeeper
Software
Site Reliability Engineer responsible for improving platform reliability, observability, and security. Duties include expanding PagerDuty on-call operations, implementing Datadog metrics and tracing, and managing credentials, certificates, encryption keys, and their rotation. The role supports infrastructure handling protected health information and offers mentorship toward broader SRE ownership.
Top Skills:
Aws AcmAws KmsAws Secrets ManagerDatadogHashicorp VaultPagerduty
Reposted One Month AgoSaved
Artificial Intelligence • Cloud • Fintech • Information Technology • Insurance • Financial Services • Big Data Analytics
Lead application readiness and remediation coordination for AWS, EOL, and vulnerability patches. Validate impacts, define smoke and regression tests, drive automation, resolve dependencies, escalate blockers, and secure production sign-off to ensure audit-ready closure.
Top Skills:
AmiApi TestingAWSCertificatesCi/Cd PipelinesContainerizationDastDatabasesDockerEc2EksLibrariesMiddlewareNew RelicNew Relic MonitorsObservability ToolingRegression AutomationRuntimesSastScaService DashboardsSmoke TestingSyntheticsTerraform
New
Track Smarter, Apply Better.
Ditch the spreadsheets. Organize your job search with our freeApplication Tracker.
Use For Free
Financial Services
Supports reliability and operations for hybrid on-premises and multi-cloud infrastructure across Azure, AWS, and GCP. Monitors production systems, responds to incidents, maintains compute, storage, networking, middleware, and identity services, and contributes to Infrastructure-as-Code, DevSecOps, GitOps, CI/CD, cloud security, observability, documentation, and post-incident reviews. Participates in on-call rotations and helps improve platform performance, resilience, and consistency.
Top Skills:
AnsibleAWSAzureBashCi/CdDatadogDevsecopsDhcpDnsDockerFirewallsGCPGitopsGrafanaKubernetesLinuxPowershellPrometheusPythonTcp/IpTerraformVlansWindows Server
Fintech • Information Technology • Software • Financial Services
The role involves designing and automating infrastructure management, improving reliability, building internal tools, and contributing to architectural decisions. Responsibilities include working with Kubernetes and managing large-scale infrastructure, while participating in on-call rotations to prevent incidents.
Top Skills:
CloudwatchDatadogDockerEfkElkGoJavaScriptKubernetesPythonTerraform
Mobile • Software
Site Reliability Engineers will work on production infrastructure, focusing on AWS and Kubernetes while ensuring high availability and customer satisfaction.
Top Skills:
AirflowAWSCircleCICloudwatchEksGrafanaMongoDBPagerdutyPingdomRustScala SparkTerraformTypescript
Reposted One Month AgoSaved
Easy Apply
Easy Apply
Big Data • Cloud • Software • Database
The Senior Site Reliability Engineer will develop and support distributed storage services, ensuring reliability and operational safety, with a focus on automation and efficiency.
Top Skills:
AWSAzureDnsGoGoogle Cloud PlatformKubernetesLinuxPythonTcp/IpTls
Big Data • Cloud • Software • Database
Seeking a Site Reliability Engineer with expertise in networking and distributed systems for building secure multi-cloud infrastructure. Responsibilities include maintaining network architecture and ensuring reliable service-to-service communication, involving a 24/7 on-call rotation.
Top Skills:
AWSAzureBgpDnsGCPIpv6KubernetesLoad BalancingMtlsService MeshTcp/IpTlsVpcsVpns
Reposted 25 Days AgoSaved
Software • Defense
Work as an SRE embedded with product teams to improve reliability by fixing application code (primarily TypeScript), building observability (Prometheus, Loki, Grafana, Alloy), defining SLIs/SLOs, leading incident response and postmortems, automating toil, and supporting deployments across on‑prem DoD and AWS environments.
Top Skills:
AlloyAWSBashContainersDockerGithub ActionsGitlab Ci/CdGoGrafanaJenkinsKubectlKubernetesLokiNode.jsPrometheusPythonTypescript
Productivity
Build and operate reliable, scalable infrastructure and backend systems. Define and monitor SLOs and SLAs, manage on-call operations, improve observability, performance, security, developer experience, and technical processes. Contribute to Temporal-based workflow orchestration, advise backend projects, manage AWS infrastructure, and support system scalability as the customer base grows.
Top Skills:
Amazon EksAWSCi/CdDockerHelmInfrastructure As CodeKubernetesLinuxTemporal
AdTech • Cloud • Digital Media • Information Technology • News + Entertainment • App development
Lead SRE for cloud-based live linear playout systems driving reliability, observability, incident response, SLIs/SLOs, automation, capacity planning, runbooks, and L1/L2 on-call support to ensure resilient distribution across NBCUniversal channels.
Top Skills:
AmagiAWSCmafDockerEsamGrafanaH.264HarrisHevcHlsImagineIp NetworkingKubernetesLinuxMicrosoft TeamsRistScte-224Scte-35ServicenowSlackSnellSplunkSrtTs
Software
Own the reliability, performance, availability, scalability, and cost efficiency of production Aurora MySQL databases on AWS. Responsibilities include observability, alerting, query optimization, replication, backups, disaster recovery, schema migrations, reader topology, cost optimization, runbooks, incident response, and database performance reviews. The role partners with engineering teams, supports HIPAA-compliant handling of PHI, and establishes reliable database operating practices.
Top Skills:
Amazon RdsAmazon RedshiftAurora MysqlAWSBashDatabricksDatadogGrafanaMySQLPercona Monitoring And Management (Pmm)Percona ToolkitPerformance InsightsPHPPrometheusPythonSnowflakeTerraform
Cloud • Software • Analytics
Develop and operate Arista’s FedRAMP CloudVision SaaS platform at scale. Responsibilities include improving reliability, scalability, observability, autoscaling, disaster recovery, capacity planning, CI/CD, network architecture, cost optimization, and cloud application security. The role develops and manages Kubernetes-native services, distributed databases, and automation using technologies such as GCP, GKE, Go, Python, Ansible, Pulumi, and Bash. Participation in a FedRAMP on-call rotation is required.
Top Skills:
AnsibleBashCi/CdDistributed DatabasesFedrampGoGoogle Cloud Platform (Gcp)Google Kubernetes Engine (Gke)KubernetesPulumiPythonSaaS
Information Technology • Software • Financial Services • Quantitative Trading
The Site Reliability Engineer will provide support and diagnose issues within a real-time, distributed environment, focusing on large-scale application and infrastructure management, with basic required skills in UNIX/Linux, networking, SQL, and scripting languages.
Top Skills:
BashPythonSQLTcp/IpUdpUnix/Linux
Fitness • Healthtech • Software
Own reliability and security of CI/CD and production services: define SLI/SLOs, lead incident response and postmortems, build observability (Datadog), operate Kubernetes and IaC (Terraform), harden pipelines with SAST/DAST/SCA and policy-as-code, and coach teams on reliability and operational best practices.
Top Skills:
AWSCi/CdConftestDastDatadogDockerGithub ActionsGoInfrastructure As CodeKubernetesKyvernoOpa/RegoPythonSastScaTerraformTypescript
eCommerce • Manufacturing
Manage Azure infrastructure and AKS clusters, build GitHub Actions CI/CD pipelines, and improve Grafana-based observability and incident response. Define SLIs, SLOs, and error budgets; maintain infrastructure as code with Pulumi; troubleshoot reliability issues; perform capacity planning and performance tuning; participate in on-call support; and document operational procedures. Collaborate with development teams to deliver scalable, reliable production systems.
Top Skills:
Azure Kubernetes Service (Aks)BashDnsGithub ActionsGrafanaIstioKubernetesLoad BalancingLokiAzurePrometheusPulumiPython
Let Your Resume Do The Work
Upload your resume to be matched with jobs you're a great fit for.
Success! We'll use this to further personalize your experience.
Top NYC, NY Companies Hiring Reliability Engineers
See AllPopular NYC, NY Engineering Job Searches
Engineering Jobs in NYC, NY
.NET Developer Jobs in NYC, NY
Android Developer Jobs in NYC, NY
Associate Software Engineer Jobs in NYC, NY
Automation Engineer Jobs in NYC, NY
AWS Engineer Jobs in NYC, NY
Backend Engineer Jobs in NYC, NY
C# Jobs in NYC, NY
C++ Jobs in NYC, NY
Cloud Engineer Jobs in NYC, NY
Controls Engineer Jobs in NYC, NY
CTO Jobs in NYC, NY
Design Engineer Jobs in NYC, NY
DevOps Engineer Jobs in NYC, NY
DevOps Jobs in NYC, NY
Director of Engineering Jobs in NYC, NY
Electrical Engineering Jobs in NYC, NY
Embedded Software Engineer Jobs in NYC, NY
Engineering Manager Jobs in NYC, NY
Field Engineer Jobs in NYC, NY
Front End Developer Jobs in NYC, NY
Full-Stack Engineer Jobs in NYC, NY
Golang Jobs in NYC, NY
Hardware Engineer Jobs in NYC, NY
Infrastructure Engineer Jobs in NYC, NY
iOS Developer Jobs in NYC, NY
Java Developer Jobs in NYC, NY
Java Full-Stack Engineer Jobs in NYC, NY
Javascript Jobs in NYC, NY
Lead Software Engineer Jobs in NYC, NY
Linux Jobs in NYC, NY
Manufacturing Engineer Jobs in NYC, NY
Mechanical Design Engineer Jobs in NYC, NY
Mechanical Engineering Jobs in NYC, NY
Mechatronics Engineering Jobs in NYC, NY
Network Engineer Jobs in NYC, NY
PHP Developer Jobs in NYC, NY
Platform Engineer Jobs in NYC, NY
Principal Architect Jobs in NYC, NY
Principal Engineer Jobs in NYC, NY
Principal Software Engineer Jobs in NYC, NY
Process Engineer Jobs in NYC, NY
Project Engineer Jobs in NYC, NY
Python Jobs in NYC, NY
QA Engineer Jobs in NYC, NY
QA Jobs in NYC, NY
Quality Assurance Automation Engineer Jobs in NYC, NY
Reliability Engineer Jobs in NYC, NY
Robotics Engineer Jobs in NYC, NY
Ruby Jobs in NYC, NY
Salesforce Developer Jobs in NYC, NY
Scala Jobs in NYC, NY
Security Engineer Jobs in NYC, NY
Software Engineer Jobs in NYC, NY
Solutions Architect Jobs in NYC, NY
Solutions Engineer Jobs in NYC, NY
SRE Engineer Jobs in NYC, NY
Staff Engineer Jobs in NYC, NY
Staff Software Engineer Jobs in NYC, NY
Systems Engineer Jobs in NYC, NY
Vice President Of Engineering Jobs in NYC, NY
Web Developer Jobs in NYC, NY
All Filters
Total selected ()
No Results
No Results
















.png)














