Maximum of 25 job preferences reached.
Top SRE Engineer Jobs in NYC, NY
Artificial Intelligence • Information Technology
Owns end-to-end reliability for Tinker, including CI/CD, observability, service-level objectives, incident response, multi-tenant isolation, resource scheduling, and vulnerability remediation. The role designs monitoring for distributed training systems, improves recovery and checkpointing, and operates Kubernetes clusters supporting heterogeneous GPU workloads. It collaborates closely with platform, security, engineering, and research teams to improve production resilience and utilization.
Top Skills:
Ci/CdCloud InfrastructureDistributed SystemsGpu WorkloadsKubernetesLoraProduction Observability
Artificial Intelligence • Marketing Tech • Software
Lead SRE to provide technical leadership for a multi-cloud, multi-region active-active content platform. Define automation and IaC strategy, design core platform and logging/observability systems, establish capacity and performance frameworks, lead cross-team reliability initiatives, improve on-call and incident response, mentor engineers, and proactively address systemic reliability and scaling challenges.
Top Skills:
Apache KafkaApache PulsarAWSCassandraChefEksGCPGkeGoGrafana AlloyGrafana LokiKubernetesLinuxNode.jsPrometheusPythonRubyScylladbShell ScriptingTempoTerraformThanos
Reposted One Month AgoSaved
Easy Apply
Easy Apply
Big Data • Cloud • Software • Database
Maintain and improve multi-cloud Kubernetes infrastructure, CI/CD (Argo Workflows/ArgoCD), observability, and networking. Build reliable continuous deployment tooling and onboarding flows, provide internal support, collaborate across Platform Engineering, contribute upstream (open-source/operators), and participate in a 24/7 on-call rotation to resolve deployment infrastructure issues.
Top Skills:
AlertingArgo WorkflowsArgocdAWSAzureCi/CdContainersDnsGCPGoKubernetesLinuxLoad BalancerObservabilityPythonService MeshTcp/IpTls
Edtech • Machine Learning • Mobile • Other • Software
Designs, operates, and improves Duolingo’s large-scale distributed systems and core infrastructure. Responsibilities include diagnosing production issues, developing platforms and automation, conducting launch reviews and root cause analyses, maintaining incident response and postmortem practices, and improving reliability, scalability, and engineering velocity. The role partners with product and platform engineering teams to reduce operational toil and ensure high-quality service delivery.
Top Skills:
DockerDynamoGoJavaKotlinKubernetesMesosMySQLNomadPostgresPython
Artificial Intelligence • Natural Language Processing • Generative AI
Own production safeguards infrastructure for Claude model launches and safety classifier deployments. Configure and verify safeguards across first-party, AWS Bedrock, and GCP Vertex platforms; lead canary rollouts, post-deployment validation, incident response, and rollback decisions. Build automation, continuous validation, repeatable deployment pipelines, and a provenance registry to reduce operational toil and configuration drift. Participate in on-call rotations and launch readiness processes.
Top Skills:
AWSGCPLlm Inference SystemsMl InfrastructurePythonRustTransformer-Based Models
Healthtech • Software
Operate and maintain AWS-hosted MERN applications and large-scale data workflows. Manage serverless and Spark-based pipelines, perform incident response and on-call duties, engineer automation to eliminate operational toil, ensure HIPAA/SOC2/HITRUST compliance, build observability and lead blameless post-mortems.
Top Skills:
Amazon EcsAmazon EksAmazon EmrAthenaAws GlueAws LambdaAws SnsAws SqsCloudwatchEc2IamJavaScriptMernMySQLNode.jsOpentofuPysparkPythonRabbitMQTerraformTypescriptVpc
Cloud • Computer Vision • Information Technology • Sales • Security • Cybersecurity
Senior SRE owning availability, automation, and observability for CI/CD platform services. Build and operate infrastructure, run on-call, lead incident response, mentor engineers, drive design/capacity planning, integrate AI-assisted workflows, and improve cross-team reliability.
Top Skills:
Active DirectoryAnsibleApache AirflowSparkAWSAzureBashBazelBitbucketCassandraChefDatadogDnsFirewall RulesGCPGitGithub ActionsGitlabGitlab CiGoGrafanaHoneycombHumio/LogscaleJenkinsKafkaKubernetesLoad BalancersMongoDBMySQLNasNew RelicNfsObject StorageOpensearchOraclePostgresPowershellPrometheusPulsarPuppetPythonRabbitMQRedis/ValkeyRedpandaRoutingSaltSanSplunkTerraformVarnishVipsWindows Server
Cloud • Security • Software • Cybersecurity
Deploy and operate scalable, highly available cloud systems; improve application, network, and infrastructure reliability, security, performance, and capacity; automate cloud deployments; monitor services and SLAs; analyze logs and events; troubleshoot infrastructure issues; and provide operational recommendations for large-scale cloud platforms.
Top Skills:
Amazon Ec2Amazon Web Services (Aws)AnsibleAws CloudwatchBashCi/CdContainerizationDockerGitGitlabHelmInfrastructure As Code (Iac)JenkinsKubernetesLog AnalysisMavenAzureNagiosOrchestrationPythonSplunkSvnVersion ControlVmware Vsphere
Cloud • Security • Software • Cybersecurity
Deploys and operates scalable, highly available cloud systems; improves application and network security, stability, speed, and capacity; automates cloud deployments; monitors and troubleshoots services to meet SLAs; analyzes logs and events; manages large-scale cloud infrastructure; and resolves infrastructure issues.
Top Skills:
Amazon Ec2Amazon EventbridgeAmazon Route 53Amazon S3Amazon Web ServicesAnsibleAws LambdaBashCi/CdDockerElastic Load BalancingJenkinsJinjaKubernetesPackerPulumiPython
Insurance
Designs and operates reliable hybrid application platforms, leading CI/CD, infrastructure automation, cloud resource management, monitoring, security scanning, container orchestration, and database self-service tooling. The role owns the enterprise CI/CD technology stack and hosting strategy across on-premises, hybrid, and cloud environments. It requires extensive collaboration with engineering, infrastructure, security, application, QA, and governance teams, along with 24/7 mission-critical support and technical leadership.
Top Skills:
AnsibleApache CamelAWSCi/CdCloudFormationConfluenceDatabase SystemsDockerGitGithub ActionsGitlab CiGradleHelmInfrastructure As CodeJenkinsJIRAKubernetesLinuxMavenMonitoring And AlertingNetworkingPythonSonatypeTerraformVulnerability Scanning
Cloud • Security
Own reliability, availability, performance, and capacity for production SaaS services across Azure, AWS, and a FedRAMP High environment. Build observability, SLOs, monitoring, infrastructure automation, and remediation workflows; lead incident response, postmortems, support escalations, and disaster recovery efforts. Manage Terraform, CI/CD, Kubernetes, WAF, networking, and observability costs while improving on-call operations and collaborating with Security, Product, Support, and Development.
Top Skills:
AksAWSAzureAzure App ServiceAzure DevopsAzure Front DoorAzure MonitorAzure Service BusAzure SqlAzure StorageAzure WafCloudflareCloudwatch Logs InsightsConsulDatadogElk StackFedrampImpervaIso 27001JenkinsJira Service ManagementJSONKubernetesMicrosoft Entra IdNist 800-53OidcPagerdutyPciPowershellPythonRedisS3SaltstackSAMLSoc 2TerraformYaml
Cloud • Information Technology • Security • Virtual Reality • Cybersecurity
Support reliability, security, scalability, and performance of the TAK CI/CD pipeline and critical infrastructure. Monitor system health, respond to incidents, patch systems, optimize performance and costs, improve DevOps practices, automate operations, and maintain technical documentation. Requires systems and network administration experience, Linux and cloud expertise, containerization, CI/CD knowledge, Zero Trust implementation experience, Security+ certification, and an active T3 investigation.
Top Skills:
AWSCi/Cd PipelinesContainerizationLinuxZero Trust
New
Cut your apply time in half.
Use ourAI Assistantto automatically fill your job applications.
Use For Free
Edtech • HR Tech • Software
Lead reliability engineering for critical SaaS services by defining SLOs, managing major incidents, improving observability, architecting scalable infrastructure, advancing deployment safety, and mentoring engineers. The role partners with engineering and product leaders to balance feature delivery with system resilience, while driving reliability standards, automation, and recruiting across the SRE organization.
Top Skills:
AWSContainer OrchestrationInfrastructure As Code
Artificial Intelligence • Healthtech • Software • Telehealth
Designs, deploys, and maintains resilient AWS and Kubernetes infrastructure. Builds automation, GitHub Actions components, internal AI-assisted operational tools, and observability systems. Leads incident response, postmortems, and SLO/SLI management while ensuring HIPAA compliance and high availability. Collaborates across teams on architecture reviews, risk reduction, clinical safety, and reliability best practices.
Top Skills:
Ai-Assisted OperationsAmazon Ec2Amazon EksAmazon RdsAmazon S3AWSBashDatadogGithub ActionsGoHelmKubernetesPythonTerraform
Fintech • Payments • Social Impact • Software
Build and manage backend applications and platform components for Chariot’s identity and payment network. Design scalable APIs and backend systems, establish engineering standards, improve monitoring and alerting, and collaborate directly with founders, product managers, and engineers. Drive projects from inception through delivery within an Agile/Scrum environment while contributing to architecture, technology strategy, and product decisions.
Top Skills:
AWSDockerGoGrpcKubernetesNode.jsPostgresRest ApisTerraform
Artificial Intelligence • Software • Conversational AI • Automation
Build and improve Replicant’s AI-native platform, including site reliability, CI/CD, developer tooling, observability, incident management, cloud infrastructure, and autonomous-agent harnesses. Own production infrastructure reliability at scale, reduce operational toil, improve deployment workflows, participate in on-call rotation, and shape platform engineering patterns. The role uses TypeScript/Node.js, Python, Terraform, Kubernetes, Helm, GCP, and modern monitoring tools.
Top Skills:
ClaudeCursorDatadogFreeswitchGCPGitlab CiGrafanaHelmKubernetesLlmsNode.jsPrometheusPythonSipTerraformTypescript
Other
Build and improve Replicant’s AI-native platform infrastructure, including site reliability, CI/CD, developer tooling, observability, incident management, and agent harness systems. Own production infrastructure patterns, deployment workflows, and developer self-service across Kubernetes and cloud environments. Participate in on-call and incident response while improving scalability, availability, and operational efficiency for real-time conversational AI traffic.
Top Skills:
ClaudeCursorDatadogFreeswitchGCPGitlab CiGrafanaHelmKubernetesLlmsNode.jsPrometheusPythonSipTerraformTypescript
Artificial Intelligence • Big Data • Information Technology • Other • Software • Database • Biotech
Provides frontline support for high-availability cloud systems, monitors infrastructure, manages incidents, coordinates outage bridges, documents technical issues, and develops automation and AI agentic workflows. The role partners with development teams to improve detection and resolution, supports AWS technologies, databases, networks, CI/CD tools, and observability platforms. This is a Monday–Thursday overnight shift within a 24/7 command center and includes some holiday coverage.
Top Skills:
Ai Agentic WorkflowsAWSAws Application SignalsBashCassandraChefCi/CdCouchdbFirewallsHarnessIntrusion Detection SystemsJavaLan/WanLoad BalancersMssqlMySQLNew RelicPowershellPrompt EngineeringProxy ServersPuppetPythonTcp/IpVirtualization
Automotive
Design and implement scalable cloud infrastructure, monitor performance, automate processes, ensure security and compliance, and lead a DevOps team.
Top Skills:
AWSBashCi/CdDockerElk StackGCPGrafanaKubernetesPrometheusPythonTerraform
eCommerce • Fintech • Payments • Software
The role involves ensuring software reliability and performance, managing incidents, developing infrastructure automation, and mentoring junior engineers within a platform team.
Top Skills:
AWSCloudFormationDatadogKubernetesOpentelemetryRubyRuby On RailsTerraform
Other
Design and operate cloud platforms supporting backend telecom services. Automate deployments, scaling, recovery, and infrastructure provisioning; monitor production systems; maintain observability, alerting, and dashboards; support incident response and on-call operations; manage CI/CD pipelines; and enable engineering, telecom, and data teams through reliable tools and infrastructure.
Top Skills:
AnsibleAWSAzureBashCassandraCircleCICloudFormationDatadogDnsDockerElasticsearchElk StackGitlab CiGoGCPGrafanaHttp/HttpsIamJaegerJenkinsKafkaKubernetesKvmLinuxNoSQLOpentelemetryPerlPrometheusPythonRubySaltstackSplunkSQLTcp/IpTerraformUnixVMware
Big Data • Cloud • Software • Database
The Site Reliability Engineer designs and builds infrastructure for a global cloud service, implements automation, and optimizes system performance while managing on-call operations.
Top Skills:
AWSDnsGCPHTTPKubernetesLinuxAzureProgramming LanguagesTls
Agency • Cloud • Professional Services • Software
Improve AWS production infrastructure reliability, observability, performance, and operational maturity. Build Terraform infrastructure, enhance CI/CD, automate operational work, manage incident response and on-call operations, lead postmortems, improve application resilience, support capacity planning and database reliability, and collaborate on security hardening and compliance. Mentor engineers and promote reliability practices across the organization.
Top Skills:
AWSCi/CdCircleCIDatadogGithub ActionsGitlab CiLinuxNew RelicPostgresRubyRuby On RailsSlisSlosTerraform
Artificial Intelligence • HR Tech • Professional Services • Software
Evaluate AI-generated documents, spreadsheets, and presentations for accuracy, rigor, domain quality, and presentation quality in incident management, reliability, and SRE. Apply specialized rubrics, identify factual and aesthetic errors, and provide structured written feedback. This is a fully remote, flexible independent-contractor engagement requiring at least five years of relevant professional experience, English fluency, and proficiency with Microsoft Office and Google Workspace.
Top Skills:
Google SlidesGoogle WorkspaceMS OfficePowerPoint
Cloud
Build and manage highly available Kubernetes platforms on AWS. Responsibilities include Kubernetes and Helm platform creation, Terraform infrastructure automation, Karpenter scaling, Istio service mesh management, CI/CD enablement, incident response, security, compliance, monitoring, cost optimization, troubleshooting, and documentation. The role supports multi-region cloud environments and requires federal-environment access eligibility, U.S. Person documentation, and occasional travel for in-person onboarding.
Top Skills:
AnsibleAWSBashCi/CdCircleCICloudFormationCloudwatchDockerEc2EcsEksElk StackGitlabGoGrafanaHelmIamIstioJenkinsKarpenterKubernetesPrometheusPythonRdsS3SpinnakerTerraformVpc
Let Your Resume Do The Work
Upload your resume to be matched with jobs you're a great fit for.
Success! We'll use this to further personalize your experience.
Top NYC, NY Companies Hiring SRE Engineers
See AllPopular NYC, NY Engineering Job Searches
Engineering Jobs in NYC, NY
.NET Developer Jobs in NYC, NY
Android Developer Jobs in NYC, NY
Associate Software Engineer Jobs in NYC, NY
Automation Engineer Jobs in NYC, NY
AWS Engineer Jobs in NYC, NY
Backend Engineer Jobs in NYC, NY
C# Jobs in NYC, NY
C++ Jobs in NYC, NY
Cloud Engineer Jobs in NYC, NY
Controls Engineer Jobs in NYC, NY
CTO Jobs in NYC, NY
Design Engineer Jobs in NYC, NY
DevOps Engineer Jobs in NYC, NY
DevOps Jobs in NYC, NY
Director of Engineering Jobs in NYC, NY
Electrical Engineering Jobs in NYC, NY
Embedded Software Engineer Jobs in NYC, NY
Engineering Manager Jobs in NYC, NY
Field Engineer Jobs in NYC, NY
Front End Developer Jobs in NYC, NY
Full-Stack Engineer Jobs in NYC, NY
Golang Jobs in NYC, NY
Hardware Engineer Jobs in NYC, NY
Infrastructure Engineer Jobs in NYC, NY
iOS Developer Jobs in NYC, NY
Java Developer Jobs in NYC, NY
Java Full-Stack Engineer Jobs in NYC, NY
Javascript Jobs in NYC, NY
Lead Software Engineer Jobs in NYC, NY
Linux Jobs in NYC, NY
Manufacturing Engineer Jobs in NYC, NY
Mechanical Design Engineer Jobs in NYC, NY
Mechanical Engineering Jobs in NYC, NY
Mechatronics Engineering Jobs in NYC, NY
Network Engineer Jobs in NYC, NY
PHP Developer Jobs in NYC, NY
Platform Engineer Jobs in NYC, NY
Principal Architect Jobs in NYC, NY
Principal Engineer Jobs in NYC, NY
Principal Software Engineer Jobs in NYC, NY
Process Engineer Jobs in NYC, NY
Project Engineer Jobs in NYC, NY
Python Jobs in NYC, NY
QA Engineer Jobs in NYC, NY
QA Jobs in NYC, NY
Quality Assurance Automation Engineer Jobs in NYC, NY
Reliability Engineer Jobs in NYC, NY
Robotics Engineer Jobs in NYC, NY
Ruby Jobs in NYC, NY
Salesforce Developer Jobs in NYC, NY
Scala Jobs in NYC, NY
Security Engineer Jobs in NYC, NY
Software Engineer Jobs in NYC, NY
Solutions Architect Jobs in NYC, NY
Solutions Engineer Jobs in NYC, NY
SRE Engineer Jobs in NYC, NY
Staff Engineer Jobs in NYC, NY
Staff Software Engineer Jobs in NYC, NY
Systems Engineer Jobs in NYC, NY
Vice President Of Engineering Jobs in NYC, NY
Web Developer Jobs in NYC, NY
All Filters
Total selected ()
No Results
No Results
















.png)

.png)















