EverOps Logo

EverOps

Lead Observability Engineer

Posted 11 Days Ago
Remote
Hiring Remotely in USA
Senior level
Remote
Hiring Remotely in USA
Senior level
Lead large-scale observability assessments and migrations across metrics, logs, and traces. Analyze telemetry volume, cardinality, retention, performance, and costs; design AWS-native OpenTelemetry architectures; build TCO models; and lead platform migration execution. Rebuild pipelines, dashboards, alerts, retention policies, ownership tagging, and instrumentation standards. Serve as a player-coach and primary technical contact for customer engineering leadership while delivering architecture, migration, commercial, and executive recommendations.
The summary above was generated by AI
Lead Observability Engineer (Remote)Overview

Some of the world’s most innovative global software and technology companies struggle to find engineering partners capable of stepping into complex environments and immediately driving meaningful outcomes. These teams need more than additional hands—they need senior engineers who can quickly understand an environment, identify the path forward, and execute without constant direction.

Enter EverOps – the premier Embedded Service Provider. We partner directly with customer engineering teams to assess and address mission-critical infrastructure, cloud, and delivery challenges.

The Challenge

EverOps is looking for a Lead Observability Engineer with deep hands-on experience across metrics, logs, and traces at very large scale to lead an observability maturity assessment, and the platform consolidation that follows it, for a high-scale consumer mobile platform.

The current estate spans multiple commercial observability vendors alongside self-managed Prometheus, Thanos, Grafana, Vector, and ELK. It carries tens of millions of active time series, tens of thousands of scrape targets, well over a thousand dashboards, and more than ten thousand alert definitions. Much of the log and trace data is sampled or dropped for cost reasons, ownership tagging is sparse, and the self-managed components carry an operational load that crowds out improvement.

The direction under evaluation is consolidation onto AWS-native observability services built on OpenTelemetry. This role requires someone who knows these systems well enough to price that move honestly, prove or disprove performance parity, and then lead the migration if the answer is go.

The Mission

As a Lead Observability Engineer, you will join our U.S.-Based Virtual Operating Center and lead an embedded TechPod as its Pod Leader: a player-coach who sets technical direction, serves as the primary point of contact for the customer’s engineering leadership, and stays hands-on in the work.

Your immediate priority is leading a two-month Observability Maturity Assessment. You’ll validate telemetry volumes, retention, sampling, and cardinality; build a total cost of ownership model; design an AWS-native target architecture; and deliver a migration plan and commercial recommendation that leadership can make a go/no-go decision on.

As discovery closes, your focus shifts to leading the platform migration: standing up the target collection pipeline and storage tiers, porting log transforms to OpenTelemetry, rebuilding dashboards and alerts, and establishing policy-driven retention and ownership tagging that make observability both less expensive and more useful.

The customer’s observability team is capable but stretched thin. You will be expected to add capacity rather than consume it: pull the data yourself, ask targeted questions, keep recommendations unbiased, and leave the team with fewer systems to run and better telemetry than they have today.

What You’ll Do
  • Estate Assessment: Build a complete picture of the observability estate across vendors, agents, collectors, query surfaces, data volumes, and operating model, grounded in live measurement rather than questionnaires alone.

  • Telemetry Analysis: Validate metric cardinality, active series, scrape target health, log volumes, trace sampling, and retention, and identify where fidelity is being lost today and why.

  • Cost Modeling: Build a total cost of ownership comparison between the current state and two to three costed target states, covering vendor spend, self-managed infrastructure, data transfer and egress, and AWS Pricing Calculator estimates.

  • Target Architecture: Design the AWS-native target across collection (ADOT / OpenTelemetry Collector), pipeline (Amazon Data Firehose), storage tiering (CloudWatch, S3, Amazon Managed Service for Prometheus), query surfaces (Amazon Managed Grafana, CloudWatch, Athena, OpenSearch), and alerting.

  • Retention & Data Classification: Partner with Security, Legal, and Engineering to classify telemetry data and design policy-driven retention, including long-term compliance archives.

  • Capability & Gap Analysis: Map current capabilities to AWS-native equivalents and make clear recommendations where no equivalent exists, such as continuous profiling.

  • Performance Parity: Define and run tests that show whether the target state holds query performance, alert latency, and data fidelity against success criteria agreed with engineering leadership.

  • Migration Planning & Execution: Size and sequence the migration, then lead it, including pipeline cutover, porting log transforms to OpenTelemetry, rebuilding dashboards, and deduplicating and rebuilding alerts.

  • OpenTelemetry Standards: Define the OpenTelemetry conventions (semantic conventions, resource attributes, collector topology, sampling strategy) that application teams will adopt as instrumentation moves over.

  • Ownership & Cost Attribution: Rebuild service ownership tagging so telemetry cost can be attributed back to the teams generating it.

  • Commercial Analysis: Reconcile platform spend, analyze licensing and commit structures, and inform vendor renewal strategy with a clear, unbiased recommendation.

  • Technical Leadership: Lead the TechPod as a player-coach, run the working cadence with the customer’s observability leadership, and keep the engagement light-touch on a busy internal team.

  • Documentation & Readouts: Produce the assessment report, target architecture, cost model, migration plan, and executive summary, and present findings and tradeoffs to engineering leadership.

You Have
  • Experience: 8+ years in SRE, DevOps, Observability, or Platform Engineering, including 4+ years owning production observability platforms and prior experience in a technical lead, staff, or principal-level role.

  • Observability at Scale: Deep experience running metrics, logging, and tracing platforms at large scale (millions of active series, multiple terabytes of logs per day) and making them cheaper and more reliable over time.

  • Prometheus Ecosystem: Advanced production experience with Prometheus and a long-term storage layer such as Thanos, Cortex, or Mimir, including PromQL, recording rules, remote write, and cardinality management.

  • Commercial Platforms: Hands-on experience with Datadog or a comparable commercial platform, including how its pricing is built across hosts, custom metrics, log indexing, and APM.

  • AWS Observability: Production experience with CloudWatch, CloudWatch Logs, Container Insights, X-Ray or Application Signals, Amazon Managed Service for Prometheus, and Amazon Managed Grafana.

  • OpenTelemetry: Production experience with the OpenTelemetry Collector (or ADOT), including pipeline design, processors, tail and head sampling, and multi-backend export.

  • Log Pipelines: Experience designing and operating log pipelines with Vector, Fluent Bit, Logstash, or Firehose, including transforms, routing, and tiered storage in S3.

  • Kubernetes: Strong production experience with EKS or Kubernetes, including DaemonSet agent sizing, kube-state-metrics, and monitoring very large clusters.

  • Infrastructure as Code: Advanced proficiency with Terraform; experience with Terragrunt or Atmos is a plus.

  • Alerting & Incident Response: Experience designing SLO-based alerting, reducing alert sprawl, and integrating with incident management tooling such as PagerDuty.

  • Cost Modeling: Ability to build defensible TCO models from usage data and pricing, and to explain the assumptions behind every number.

  • Automation: Strong scripting ability using Python, Go, or Bash to pull usage data, analyze telemetry, and automate migration work.

  • Discovery & Assessment: Demonstrated ability to enter an unfamiliar environment, measure it directly, and produce a defensible current-state picture and recommendation in weeks rather than quarters.

  • Communication: Ability to explain technical, cost, and compliance tradeoffs to engineers, Security and Legal stakeholders, and executive leadership.

Extra Awesome
  • Vendor Migration: Experience migrating off Datadog, Splunk, New Relic, or similar platforms onto AWS-native or open-source observability stacks.

  • Dashboards & Alerts as Code: Experience automating large-scale dashboard and alert migrations using Grafana provisioning, Grafonnet, Terraform providers, or conversion tooling.

  • Grafana Ecosystem: Experience with Grafana Alloy, Beyla, or other eBPF-based instrumentation.

  • Continuous Profiling: Experience with Pyroscope, Parca, or comparable profiling tools.

  • Data Governance: Experience designing telemetry retention and data handling to meet SOC 2, GDPR, privacy, or similar compliance requirements.

  • Analytics Platforms: Familiarity with Athena, OpenSearch, Databricks, or similar platforms used for log analytics and long-term telemetry queries.

  • Consumer Scale: Experience with high-traffic B2C platforms where telemetry volume tracks tens of millions of users.

  • Consulting: Experience leading assessment or discovery engagements that end in a committed, funded program of work.

  • Certifications: Prometheus Certified Associate, OpenTelemetry Certified Associate, AWS Certified DevOps Engineer – Professional, AWS Certified Solutions Architect – Professional, CKA, or similar.

Benefits
  • 100% Remote Workplace: We’ve been remote since Day 1!

  • Unlimited Paid Time Off.

  • Equity: Become a true owner of the company.

  • 401K with company contribution and sponsored healthcare.

  • Professional Growth: Access to training and certification programs to accelerate your career.

Similar Jobs

One Month Ago
In-Office or Remote
Home, TN, USA
107K-284K Annually
Senior level
107K-284K Annually
Senior level
Fitness • Healthtech • Retail • Pharmaceutical
Lead design, build, and operate a scalable observability platform for logs, metrics, and traces across cloud and datacenters. Drive OpenTelemetry adoption, architect high-throughput telemetry pipelines, implement SLOs and platform health checks, automate operations, participate in on-call rotations, and provide technical leadership and mentorship across engineering teams.
Top Skills: Argo CdCloudFormationDockerGoGrafanaHelmJavaKubernetesKustomizeLokiMimirMySQLOpentelemetryOtlpPostgresTempoTerraform
4 Days Ago
Remote
IL, USA
100K-171K Annually
Senior level
100K-171K Annually
Senior level
Insurance
Lead design and build of AI-powered observability and AIOps platforms: agentic AI agents, predictive insights, automated remediation, integrations with observability tools, developer-first tooling, cloud-native scalable services, and end-to-end automation across hybrid environments. Serve as technical lead, define enterprise architecture, drive CI/CD and IaC practices, mentor engineers, and collaborate with cross-functional teams to improve reliability and incident response.
Top Skills: Ai Orchestration FrameworksAiopsApi DevelopmentAppdynamicsCi/CdDatadogDynatraceEvent-Driven SystemsInfrastructure-As-CodeJavaKubernetesLinuxLlmsMicroservicesNew RelicNode.jsOpentelemetry (Otel)PythonReactSpring BootWindows
20 Minutes Ago
Remote or Hybrid
United States
93K-135K Annually
Junior
93K-135K Annually
Junior
Fintech • Information Technology • Insurance • Financial Services • Big Data Analytics
Defines customer needs, personas, problem statements, value propositions, and product solutions for disability and absence insurance. Analyzes data and market trends, identifies gaps, leads process reengineering using AI and APIs, manages the product backlog, and partners with product, operations, technology, customers, and stakeholders to improve outcomes and efficiency.
Top Skills: AgileAIAPIsData Integration

What you need to know about the NYC Tech Scene

As the undisputed financial capital of the world, New York City is an epicenter of startup funding activity. The city has a thriving fintech scene and is a major player in verticals ranging from AI to biotech, cybersecurity and digital media. It also has universities like NYU, Columbia and Cornell Tech attracting students and researchers from across the globe, providing the ecosystem with a constant influx of world-class talent. And its East Coast location and three international airports make it a perfect spot for European companies establishing a foothold in the United States.

Key Facts About NYC Tech

  • Number of Tech Workers: 549,200; 6% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Capgemini, Bloomberg, IBM, Spotify
  • Key Industries: Artificial intelligence, Fintech
  • Funding Landscape: $25.5 billion in venture capital funding in 2024 (Pitchbook)
  • Notable Investors: Greycroft, Thrive Capital, Union Square Ventures, FirstMark Capital, Tiger Global Management, Tribeca Venture Partners, Insight Partners, Two Sigma Ventures
  • Research Centers and Universities: Columbia University, New York University, Fordham University, CUNY, AI Now Institute, Flatiron Institute, C.N. Yang Institute for Theoretical Physics, NASA Space Radiation Laboratory

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account