Top SRE Engineer Jobs in NYC, NY

5 Days AgoSaved
In-Office
New York, NY
350K-475K Annually
Entry level
350K-475K Annually
Entry level
Artificial Intelligence • Information Technology
Owns end-to-end reliability for Tinker, including CI/CD, observability, service-level objectives, incident response, multi-tenant isolation, resource scheduling, and vulnerability remediation. The role designs monitoring for distributed training systems, improves recovery and checkpointing, and operates Kubernetes clusters supporting heterogeneous GPU workloads. It collaborates closely with platform, security, engineering, and research teams to improve production resilience and utilization.
Top Skills: Ci/CdCloud InfrastructureDistributed SystemsGpu WorkloadsKubernetesLoraProduction Observability
Reposted One Month AgoSaved
Easy Apply
Hybrid
New York, NY
Easy Apply
184K-240K Annually
Senior level
184K-240K Annually
Senior level
Artificial Intelligence • Marketing Tech • Software
Lead SRE to provide technical leadership for a multi-cloud, multi-region active-active content platform. Define automation and IaC strategy, design core platform and logging/observability systems, establish capacity and performance frameworks, lead cross-team reliability initiatives, improve on-call and incident response, mentor engineers, and proactively address systemic reliability and scaling challenges.
Top Skills: Apache KafkaApache PulsarAWSCassandraChefEksGCPGkeGoGrafana AlloyGrafana LokiKubernetesLinuxNode.jsPrometheusPythonRubyScylladbShell ScriptingTempoTerraformThanos
Reposted One Month AgoSaved
Easy Apply
Remote or Hybrid
New York, NY
Easy Apply
127K-249K Annually
Senior level
127K-249K Annually
Senior level
Big Data • Cloud • Software • Database
Maintain and improve multi-cloud Kubernetes infrastructure, CI/CD (Argo Workflows/ArgoCD), observability, and networking. Build reliable continuous deployment tooling and onboarding flows, provide internal support, collaborate across Platform Engineering, contribute upstream (open-source/operators), and participate in a 24/7 on-call rotation to resolve deployment infrastructure issues.
Top Skills: AlertingArgo WorkflowsArgocdAWSAzureCi/CdContainersDnsGCPGoKubernetesLinuxLoad BalancerObservabilityPythonService MeshTcp/IpTls
8 Days AgoSaved
Hybrid
New York, NY
183K-247K Annually
Senior level
183K-247K Annually
Senior level
Edtech • Machine Learning • Mobile • Other • Software
Designs, operates, and improves Duolingo’s large-scale distributed systems and core infrastructure. Responsibilities include diagnosing production issues, developing platforms and automation, conducting launch reviews and root cause analyses, maintaining incident response and postmortem practices, and improving reliability, scalability, and engineering velocity. The role partners with product and platform engineering teams to reduce operational toil and ensure high-quality service delivery.
Top Skills: DockerDynamoGoJavaKotlinKubernetesMesosMySQLNomadPostgresPython
8 Days AgoSaved
In-Office or Remote
New York, NY
320K-485K Annually
Senior level
320K-485K Annually
Senior level
Artificial Intelligence • Natural Language Processing • Generative AI
Own production safeguards infrastructure for Claude model launches and safety classifier deployments. Configure and verify safeguards across first-party, AWS Bedrock, and GCP Vertex platforms; lead canary rollouts, post-deployment validation, incident response, and rollback decisions. Build automation, continuous validation, repeatable deployment pipelines, and a provenance registry to reduce operational toil and configuration drift. Participate in on-call rotations and launch readiness processes.
Top Skills: AWSGCPLlm Inference SystemsMl InfrastructurePythonRustTransformer-Based Models
Reposted 27 Days AgoSaved
Easy Apply
Remote
New York, NY
Easy Apply
100K-110K Annually
Mid level
100K-110K Annually
Mid level
Healthtech • Software
Operate and maintain AWS-hosted MERN applications and large-scale data workflows. Manage serverless and Spark-based pipelines, perform incident response and on-call duties, engineer automation to eliminate operational toil, ensure HIPAA/SOC2/HITRUST compliance, build observability and lead blameless post-mortems.
Top Skills: Amazon EcsAmazon EksAmazon EmrAthenaAws GlueAws LambdaAws SnsAws SqsCloudwatchEc2IamJavaScriptMernMySQLNode.jsOpentofuPysparkPythonRabbitMQTerraformTypescriptVpc
28 Days AgoSaved
Remote or Hybrid
New York, NY
140K-215K Annually
Senior level
140K-215K Annually
Senior level
Cloud • Computer Vision • Information Technology • Sales • Security • Cybersecurity
Senior SRE owning availability, automation, and observability for CI/CD platform services. Build and operate infrastructure, run on-call, lead incident response, mentor engineers, drive design/capacity planning, integrate AI-assisted workflows, and improve cross-team reliability.
Top Skills: Active DirectoryAnsibleApache AirflowSparkAWSAzureBashBazelBitbucketCassandraChefDatadogDnsFirewall RulesGCPGitGithub ActionsGitlabGitlab CiGoGrafanaHoneycombHumio/LogscaleJenkinsKafkaKubernetesLoad BalancersMongoDBMySQLNasNew RelicNfsObject StorageOpensearchOraclePostgresPowershellPrometheusPulsarPuppetPythonRabbitMQRedis/ValkeyRedpandaRoutingSaltSanSplunkTerraformVarnishVipsWindows Server
YesterdaySaved
In-Office or Remote
New York, NY
186K-219K Annually
Senior level
186K-219K Annually
Senior level
Cloud • Security • Software • Cybersecurity
Deploy and operate scalable, highly available cloud systems; improve application, network, and infrastructure reliability, security, performance, and capacity; automate cloud deployments; monitor services and SLAs; analyze logs and events; troubleshoot infrastructure issues; and provide operational recommendations for large-scale cloud platforms.
Top Skills: Amazon Ec2Amazon Web Services (Aws)AnsibleAws CloudwatchBashCi/CdContainerizationDockerGitGitlabHelmInfrastructure As Code (Iac)JenkinsKubernetesLog AnalysisMavenAzureNagiosOrchestrationPythonSplunkSvnVersion ControlVmware Vsphere
YesterdaySaved
In-Office or Remote
New York, NY
138K-171K Annually
Junior
138K-171K Annually
Junior
Cloud • Security • Software • Cybersecurity
Deploys and operates scalable, highly available cloud systems; improves application and network security, stability, speed, and capacity; automates cloud deployments; monitors and troubleshoots services to meet SLAs; analyzes logs and events; manages large-scale cloud infrastructure; and resolves infrastructure issues.
Top Skills: Amazon Ec2Amazon EventbridgeAmazon Route 53Amazon S3Amazon Web ServicesAnsibleAws LambdaBashCi/CdDockerElastic Load BalancingJenkinsJinjaKubernetesPackerPulumiPython
YesterdaySaved
Remote
New York, NY
105K-252K Annually
Expert/Leader
105K-252K Annually
Expert/Leader
Insurance
Designs and operates reliable hybrid application platforms, leading CI/CD, infrastructure automation, cloud resource management, monitoring, security scanning, container orchestration, and database self-service tooling. The role owns the enterprise CI/CD technology stack and hosting strategy across on-premises, hybrid, and cloud environments. It requires extensive collaboration with engineering, infrastructure, security, application, QA, and governance teams, along with 24/7 mission-critical support and technical leadership.
Top Skills: AnsibleApache CamelAWSCi/CdCloudFormationConfluenceDatabase SystemsDockerGitGithub ActionsGitlab CiGradleHelmInfrastructure As CodeJenkinsJIRAKubernetesLinuxMavenMonitoring And AlertingNetworkingPythonSonatypeTerraformVulnerability Scanning
2 Days AgoSaved
Remote
New York, NY
130K-160K Annually
Senior level
130K-160K Annually
Senior level
Cloud • Security
Own reliability, availability, performance, and capacity for production SaaS services across Azure, AWS, and a FedRAMP High environment. Build observability, SLOs, monitoring, infrastructure automation, and remediation workflows; lead incident response, postmortems, support escalations, and disaster recovery efforts. Manage Terraform, CI/CD, Kubernetes, WAF, networking, and observability costs while improving on-call operations and collaborating with Security, Product, Support, and Development.
Top Skills: AksAWSAzureAzure App ServiceAzure DevopsAzure Front DoorAzure MonitorAzure Service BusAzure SqlAzure StorageAzure WafCloudflareCloudwatch Logs InsightsConsulDatadogElk StackFedrampImpervaIso 27001JenkinsJira Service ManagementJSONKubernetesMicrosoft Entra IdNist 800-53OidcPagerdutyPciPowershellPythonRedisS3SaltstackSAMLSoc 2TerraformYaml
2 Days AgoSaved
Remote
New York, NY
Mid level
Mid level
Cloud • Information Technology • Security • Virtual Reality • Cybersecurity
Support reliability, security, scalability, and performance of the TAK CI/CD pipeline and critical infrastructure. Monitor system health, respond to incidents, patch systems, optimize performance and costs, improve DevOps practices, automate operations, and maintain technical documentation. Requires systems and network administration experience, Linux and cloud expertise, containerization, CI/CD knowledge, Zero Trust implementation experience, Security+ certification, and an active T3 investigation.
Top Skills: AWSCi/Cd PipelinesContainerizationLinuxZero Trust
New

Cut your apply time in half.

Use ourAI Assistantto automatically fill your job applications.

Use For Free
Application Tracker Preview
2 Days AgoSaved
Remote or Hybrid
New York, NY
145K-193K Annually
Senior level
145K-193K Annually
Senior level
Edtech • HR Tech • Software
Lead reliability engineering for critical SaaS services by defining SLOs, managing major incidents, improving observability, architecting scalable infrastructure, advancing deployment safety, and mentoring engineers. The role partners with engineering and product leaders to balance feature delivery with system resilience, while driving reliability standards, automation, and recruiting across the SRE organization.
Top Skills: AWSContainer OrchestrationInfrastructure As Code
12 Days AgoSaved
In-Office or Remote
New York, NY
135K-160K Annually
Senior level
135K-160K Annually
Senior level
Artificial Intelligence • Healthtech • Software • Telehealth
Designs, deploys, and maintains resilient AWS and Kubernetes infrastructure. Builds automation, GitHub Actions components, internal AI-assisted operational tools, and observability systems. Leads incident response, postmortems, and SLO/SLI management while ensuring HIPAA compliance and high availability. Collaborates across teams on architecture reviews, risk reduction, clinical safety, and reliability best practices.
Top Skills: Ai-Assisted OperationsAmazon Ec2Amazon EksAmazon RdsAmazon S3AWSBashDatadogGithub ActionsGoHelmKubernetesPythonTerraform
12 Days AgoSaved
Hybrid
New York, NY
180K-230K Annually
Mid level
180K-230K Annually
Mid level
Fintech • Payments • Social Impact • Software
Build and manage backend applications and platform components for Chariot’s identity and payment network. Design scalable APIs and backend systems, establish engineering standards, improve monitoring and alerting, and collaborate directly with founders, product managers, and engineers. Drive projects from inception through delivery within an Agile/Scrum environment while contributing to architecture, technology strategy, and product decisions.
Top Skills: AWSDockerGoGrpcKubernetesNode.jsPostgresRest ApisTerraform
3 Days AgoSaved
Remote
New York, NY
Senior level
Senior level
Artificial Intelligence • Software • Conversational AI • Automation
Build and improve Replicant’s AI-native platform, including site reliability, CI/CD, developer tooling, observability, incident management, cloud infrastructure, and autonomous-agent harnesses. Own production infrastructure reliability at scale, reduce operational toil, improve deployment workflows, participate in on-call rotation, and shape platform engineering patterns. The role uses TypeScript/Node.js, Python, Terraform, Kubernetes, Helm, GCP, and modern monitoring tools.
Top Skills: ClaudeCursorDatadogFreeswitchGCPGitlab CiGrafanaHelmKubernetesLlmsNode.jsPrometheusPythonSipTerraformTypescript
3 Days AgoSaved
Remote
New York, NY
Senior level
Senior level
Other
Build and improve Replicant’s AI-native platform infrastructure, including site reliability, CI/CD, developer tooling, observability, incident management, and agent harness systems. Own production infrastructure patterns, deployment workflows, and developer self-service across Kubernetes and cloud environments. Participate in on-call and incident response while improving scalability, availability, and operational efficiency for real-time conversational AI traffic.
Top Skills: ClaudeCursorDatadogFreeswitchGCPGitlab CiGrafanaHelmKubernetesLlmsNode.jsPrometheusPythonSipTerraformTypescript
3 Days AgoSaved
Remote or Hybrid
New York, NY
106K-119K Hourly
Mid level
106K-119K Hourly
Mid level
Artificial Intelligence • Big Data • Information Technology • Other • Software • Database • Biotech
Provides frontline support for high-availability cloud systems, monitors infrastructure, manages incidents, coordinates outage bridges, documents technical issues, and develops automation and AI agentic workflows. The role partners with development teams to improve detection and resolution, supports AWS technologies, databases, networks, CI/CD tools, and observability platforms. This is a Monday–Thursday overnight shift within a 24/7 command center and includes some holiday coverage.
Top Skills: Ai Agentic WorkflowsAWSAws Application SignalsBashCassandraChefCi/CdCouchdbFirewallsHarnessIntrusion Detection SystemsJavaLan/WanLoad BalancersMssqlMySQLNew RelicPowershellPrompt EngineeringProxy ServersPuppetPythonTcp/IpVirtualization
Reposted 3 Days AgoSaved
Remote
New York, NY
Senior level
Senior level
Automotive
Design and implement scalable cloud infrastructure, monitor performance, automate processes, ensure security and compliance, and lead a DevOps team.
Top Skills: AWSBashCi/CdDockerElk StackGCPGrafanaKubernetesPrometheusPythonTerraform
Reposted One Month AgoSaved
Remote or Hybrid
New York, NY
175K-200K Annually
Senior level
175K-200K Annually
Senior level
eCommerce • Fintech • Payments • Software
The role involves ensuring software reliability and performance, managing incidents, developing infrastructure automation, and mentoring junior engineers within a platform team.
Top Skills: AWSCloudFormationDatadogKubernetesOpentelemetryRubyRuby On RailsTerraform
4 Days AgoSaved
Remote
New York, NY
Entry level
Entry level
Other
Design and operate cloud platforms supporting backend telecom services. Automate deployments, scaling, recovery, and infrastructure provisioning; monitor production systems; maintain observability, alerting, and dashboards; support incident response and on-call operations; manage CI/CD pipelines; and enable engineering, telecom, and data teams through reliable tools and infrastructure.
Top Skills: AnsibleAWSAzureBashCassandraCircleCICloudFormationDatadogDnsDockerElasticsearchElk StackGitlab CiGoGCPGrafanaHttp/HttpsIamJaegerJenkinsKafkaKubernetesKvmLinuxNoSQLOpentelemetryPerlPrometheusPythonRubySaltstackSplunkSQLTcp/IpTerraformUnixVMware
Reposted One Month AgoSaved
Easy Apply
Hybrid
New York, NY
Easy Apply
111K-218K Annually
Mid level
111K-218K Annually
Mid level
Big Data • Cloud • Software • Database
The Site Reliability Engineer designs and builds infrastructure for a global cloud service, implements automation, and optimizes system performance while managing on-call operations.
Top Skills: AWSDnsGCPHTTPKubernetesLinuxAzureProgramming LanguagesTls
4 Days AgoSaved
Remote
New York, NY
Senior level
Senior level
Agency • Cloud • Professional Services • Software
Improve AWS production infrastructure reliability, observability, performance, and operational maturity. Build Terraform infrastructure, enhance CI/CD, automate operational work, manage incident response and on-call operations, lead postmortems, improve application resilience, support capacity planning and database reliability, and collaborate on security hardening and compliance. Mentor engineers and promote reliability practices across the organization.
Top Skills: AWSCi/CdCircleCIDatadogGithub ActionsGitlab CiLinuxNew RelicPostgresRubyRuby On RailsSlisSlosTerraform
4 Days AgoSaved
Remote
New York, NY
80-120 Hourly
Senior level
80-120 Hourly
Senior level
Artificial Intelligence • HR Tech • Professional Services • Software
Evaluate AI-generated documents, spreadsheets, and presentations for accuracy, rigor, domain quality, and presentation quality in incident management, reliability, and SRE. Apply specialized rubrics, identify factual and aesthetic errors, and provide structured written feedback. This is a fully remote, flexible independent-contractor engagement requiring at least five years of relevant professional experience, English fluency, and proficiency with Microsoft Office and Google Workspace.
Top Skills: Google SlidesGoogle WorkspaceMS OfficePowerPoint
15 Days AgoSaved
Hybrid
New York, NY
194K-267K Annually
Senior level
194K-267K Annually
Senior level
Cloud
Build and manage highly available Kubernetes platforms on AWS. Responsibilities include Kubernetes and Helm platform creation, Terraform infrastructure automation, Karpenter scaling, Istio service mesh management, CI/CD enablement, incident response, security, compliance, monitoring, cost optimization, troubleshooting, and documentation. The role supports multi-region cloud environments and requires federal-environment access eligibility, U.S. Person documentation, and occasional travel for in-person onboarding.
Top Skills: AnsibleAWSBashCi/CdCircleCICloudFormationCloudwatchDockerEc2EcsEksElk StackGitlabGoGrafanaHelmIamIstioJenkinsKarpenterKubernetesPrometheusPythonRdsS3SpinnakerTerraformVpc
All Filters
JobType
New Jobs
Job Category
Experience
Industry
Company Name
Company Size

Sign up now Access later

Create Free Account