EverOps Logo

EverOps

Lead DevOps Engineer

Posted Yesterday
Remote
Hiring Remotely in USA
Senior level
Remote
Hiring Remotely in USA
Senior level
Lead hands-on cloud migration from legacy hosted infrastructure into AWS, resolve complex networking issues, build IaC with Terraform, modernize CentOS workloads, implement CI/CD and observability, design AWS disaster recovery, support production cutovers, and provide senior technical leadership, documentation, and runbooks to drive migrations and stabilization.
The summary above was generated by AI
Lead DevOps Engineer (Remote)

Overview

Some of the world’s most innovative global software and technology companies struggle to find engineering partners capable of stepping into complex environments and immediately driving meaningful outcomes. These teams need more than additional hands—they need senior engineers who can quickly understand an environment, identify the path forward, and execute without constant direction.

Enter EverOps – the premier Embedded Service Provider. We partner directly with customer engineering teams to assess and address mission-critical infrastructure, cloud, and delivery challenges.

The Challenge

EverOps is looking for a Lead DevOps Engineer with deep AWS infrastructure experience and unusually strong networking expertise to support a complex cloud migration and modernization initiative within a high-transaction payments environment.

You’ll be joining an active migration already in motion, where timelines are compressed, dependencies are not always fully documented, and the environment spans legacy infrastructure, AWS, application platforms, databases, networking, and production operations.

This role requires someone who can get productive quickly, work through ambiguity, and independently turn broad objectives into executable technical work.

The Mission

As a Lead DevOps Engineer, you will join our U.S.-Based Virtual Operating Center and embed directly with a customer engineering team.

Your immediate priority will be providing hands-on engineering leadership and execution for an active migration from legacy hosted infrastructure into AWS. You’ll troubleshoot migration blockers, assess network and infrastructure dependencies, build cloud infrastructure, and help move production workloads safely.

As the immediate migration effort stabilizes, your focus will expand into modernizing legacy CentOS workloads, building the supporting deployment and observability platform, and implementing AWS-based disaster recovery for a business-critical payments platform.

This is a highly autonomous role. You will be expected to identify what needs to happen, define technical deliverables, communicate risks and dependencies clearly, and drive work through completion without requiring step-by-step direction.

What You’ll Do
  • Cloud Migration: Provide hands-on engineering support for the migration of production workloads from legacy hosted infrastructure into AWS.

  • Network Engineering: Diagnose and resolve complex connectivity, routing, DNS, firewall, VPN, load-balancing, security-group, and hybrid-networking issues impacting migrations and production systems.

  • Migration Planning: Assess undocumented or partially documented environments, identify dependencies and blockers, and translate findings into practical migration plans and technical workstreams.

  • AWS Infrastructure: Design, build, troubleshoot, and improve production AWS environments using modern infrastructure-as-code practices.

  • Legacy Modernization: Help retire legacy CentOS 7 systems and migrate applications onto a modern, supportable platform.

  • Platform Engineering: Implement and improve CI/CD, monitoring, observability, secrets management, configuration management, and deployment automation.

  • Disaster Recovery: Design and implement AWS-based disaster recovery capabilities, including infrastructure, replication dependencies, recovery procedures, and operational runbooks.

  • Production Cutovers: Support migration rehearsals, rollback planning, production cutovers, validation, and post-migration stabilization.

  • Technical Ownership: Independently define and execute technical deliverables, surface risks early, and drive issues to resolution.

  • Documentation: Produce useful architecture documentation, migration plans, operational runbooks, dependency maps, and knowledge-transfer materials.

  • Technical Leadership: Serve as a senior technical partner to customer engineers and EverOps team members, providing direction when ambiguity or complex infrastructure decisions arise.

You Have
  • Experience: 7+ years of professional experience in DevOps, Cloud Engineering, SRE, Infrastructure Engineering, or a related discipline, with significant production AWS experience.

  • AWS Expertise: Deep hands-on experience designing, operating, troubleshooting, and migrating production workloads in AWS.

  • Networking Depth: Strong knowledge of TCP/IP, routing, subnetting, DNS, NAT, firewalls, VPNs, proxies, load balancers, security groups, network ACLs, and AWS networking services.

  • AWS Networking: Production experience with VPC architecture, Transit Gateway, Route 53, ALB/NLB, PrivateLink/VPC endpoints, VPN connectivity, and multi-account or hybrid-network environments.

  • Migration Experience: Proven experience executing data-center, hosted-infrastructure, lift-and-shift, re-platforming, or cloud-modernization migrations involving live production workloads.

  • Infrastructure as Code: Advanced proficiency with Terraform and experience managing production infrastructure through version-controlled IaC.

  • Linux: Strong Linux systems administration and troubleshooting skills, including experience with legacy environments and operating-system modernization.

  • Containers: Production experience with Docker and container orchestration platforms such as ECS, EKS, or Kubernetes.

  • CI/CD: Experience designing and operating modern CI/CD pipelines using tools such as GitHub Actions, Jenkins, Argo CD, or similar platforms.

  • Observability: Experience implementing and troubleshooting monitoring, logging, metrics, and alerting platforms such as Datadog, Prometheus, Grafana, or comparable tooling.

  • Automation: Strong scripting ability using Python, Bash, or similar languages to automate infrastructure and operational workflows.

  • Production Operations: Experience supporting business-critical production environments where uptime, change control, and careful migration planning matter.

  • Autonomy: Demonstrated ability to enter an unfamiliar environment, independently identify priorities, define a path forward, and execute with minimal supervision.

  • Communication: Ability to clearly explain technical risks, dependencies, tradeoffs, and recommendations to both engineers and technical leadership.

Extra Awesome
  • Fintech / Payments: Experience working in payments, financial services, banking, or another highly regulated, high-transaction environment.

  • Rackspace / Hosted Infrastructure: Experience migrating workloads out of Rackspace or similar managed hosting / colocation environments.

  • Migration Leadership: Experience owning end-to-end infrastructure migrations, including discovery, dependency mapping, rehearsal, cutover, rollback, and stabilization.

  • Disaster Recovery: Hands-on experience designing and implementing AWS DR environments, replication strategies, recovery runbooks, and RTO/RPO objectives.

  • Database Infrastructure: Familiarity with infrastructure supporting Aurora/MySQL, PostgreSQL, Cassandra, or other distributed data platforms.

  • Security: Experience operating within environments involving PCI, SOC 2, or similar security and compliance requirements.

  • GitOps: Experience with GitOps and pull-request-driven infrastructure workflows using tools such as Atlantis, Terraform Cloud/Enterprise, Scalr, or Argo CD.

  • Platform Engineering: Experience building internal platforms that standardize deployment, observability, secrets management, and infrastructure consumption.

  • Certifications: AWS Certified Solutions Architect – Professional, AWS Certified Advanced Networking – Specialty, CKA, or similar advanced certifications.

Benefits
  • 100% Remote Workplace: We’ve been remote since Day 1!

  • Unlimited Paid Time Off.

  • Equity: Become a true owner of the company.

  • 401K with company contribution and sponsored healthcare.

  • Professional Growth: Access to training and certification programs to accelerate your career.

Similar Jobs

9 Days Ago
In-Office or Remote
Basking Ridge, NJ, USA
113K-193K Annually
Senior level
113K-193K Annually
Senior level
Artificial Intelligence • Big Data • Healthtech • Information Technology • Machine Learning • Software • Analytics
Lead production operations and incident management for customer-facing applications. Coordinate releases, drive P1/P2 incident response and RCAs, execute DR exercises, manage capacity and CMDB, support cloud platform operations (Kubernetes, Kafka, GCP/HCP), and lead automation and reliability improvements with engineering and vendor partners.
Top Skills: AnsibleAWSAzureBashCmdbDatadogDockerEvent Driven ArchitectureGCPGrafanaHcp ConsoleJavaKafkaKubernetesLoad BalancerMessaging SystemsMicroservicesNosql DatabasesPrometheusPythonRelational DatabasesRest ApiTerraform
16 Days Ago
In-Office or Remote
Senior level
Senior level
Information Technology • Software
Lead design, build, and operate cloud infrastructure on GCP using Terraform and Kubernetes. Build CI/CD with GitHub Actions/Jenkins, improve observability and IAM, support hybrid legacy systems (Windows, Oracle, SQL Server, AWS), collaborate with data and engineering teams, detect infrastructure drift, and mentor engineers to drive reliability, security, and automation across environments.
Top Skills: Auth0AWSAzureBashBigQueryDagsterDbtGithub ActionsGoogle Cloud Platform (Gcp)JenkinsKubernetes (Gke)OidcOraclePythonSnowflakeSQL ServerTerraformWindows ServerWorkload Identity
2 Days Ago
Easy Apply
Remote or Hybrid
United States
Easy Apply
160K-215K Annually
Senior level
160K-215K Annually
Senior level
Fintech • News + Entertainment • Software • Database • Financial Services
Lead and mentor a DevOps team to design, implement, and maintain CI/CD pipelines, cloud infrastructure, IaC, container orchestration, monitoring, security controls, disaster recovery, and on-call support to ensure scalable, reliable, and secure systems.
Top Skills: AlbAWSCloudwatchDatadogDockerEbsEc2EcsElbGithub ActionsJenkinsKubernetesLinuxRdsRoute53S3SaltstackTerraformVpc

What you need to know about the NYC Tech Scene

As the undisputed financial capital of the world, New York City is an epicenter of startup funding activity. The city has a thriving fintech scene and is a major player in verticals ranging from AI to biotech, cybersecurity and digital media. It also has universities like NYU, Columbia and Cornell Tech attracting students and researchers from across the globe, providing the ecosystem with a constant influx of world-class talent. And its East Coast location and three international airports make it a perfect spot for European companies establishing a foothold in the United States.

Key Facts About NYC Tech

  • Number of Tech Workers: 549,200; 6% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Capgemini, Bloomberg, IBM, Spotify
  • Key Industries: Artificial intelligence, Fintech
  • Funding Landscape: $25.5 billion in venture capital funding in 2024 (Pitchbook)
  • Notable Investors: Greycroft, Thrive Capital, Union Square Ventures, FirstMark Capital, Tiger Global Management, Tribeca Venture Partners, Insight Partners, Two Sigma Ventures
  • Research Centers and Universities: Columbia University, New York University, Fordham University, CUNY, AI Now Institute, Flatiron Institute, C.N. Yang Institute for Theoretical Physics, NASA Space Radiation Laboratory

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account