Nava Benefits Logo

Nava Benefits

Data Engineer

Posted 5 Hours Ago
Remote
Hiring Remotely in USA
Entry level
Remote
Hiring Remotely in USA
Entry level
Own production data ingestion pipelines for healthcare and benefits data, including normalization, mapping, validation, reconciliation, observability, and secure handling of sensitive information. Unify census data processes, scale the platform, improve identity matching, and build automation and TypeScript tooling. The role requires strong Python and SQL skills, production orchestration experience, reliable testing and data-quality practices, and effective communication with technical and nontechnical stakeholders.
The summary above was generated by AI

Who We Are

Nava is on a mission to #fixhealthcare. Nearly 160M Americans rely on their employers for healthcare — yet the system is broken, bloated, and dominated by incumbents who resist change. Nava fuses deep benefits expertise with cutting-edge technology to deliver a modern, transparent, and affordable healthcare experience.

Founded by seasoned entrepreneurs and backed by leading investors, Nava is one of the fastest-growing benefits brokerages in the country. We’re the first to combine brokerage know-how with proprietary tech like HQ: our AI-powered benefits platform, designed to help HR teams take control of renewals, simplify strategy, eliminate spreadsheet chaos, and empower employees to navigate their benefits with ease.

In a $50B industry hungry for change, Nava is built to win — lowering costs for employers, delighting employees, and reshaping how healthcare works for millions of Americans.

About This Role & Why It Matters

We’re hiring a Data Engineer to own the external pipes of data into Nava’s domain models. “Data engineer” means many things; here it means one thing. Health and benefits data arrives from many sources in many shapes: eligibility (census) files from HR and payroll systems, carrier feeds, member IDs, enrollment elections and claims. Your job is to ingest it, normalize it to our canonical models, then map and validate it so it can be associated with everything else we know about an employer and their people. We are a concentrated data team whose job is the hygiene, quality and leverageability of external data inside the Nava ecosystem.

What makes this opportunity unique: clean, connected data is what unlocks the rest of Nava. It lets our AI help an employee understand their benefits and utilization, and choose the right plan for their family next year. It puts a member ID in someone’s hand at the point of care. And it drives a data-driven conversation between an employer and their broker, which is Nava: the optimal plan options for their employee base at renewal, and audits that check whether carriers are billing for the people the HR system actually has enrolled. That is wrong more often than you would think, and finding it matters to our clients.

What You’ll Accomplish in Your First Year

  • Own ingestion end to end: Every external source lands through pipelines that fail loudly on bad input, reconcile counts from source to normalized tables, and surface problems to us before a downstream team or a client does.

  • Unify the census platform: Our employee eligibility and elections data is a series of processes built over five years by different teams. Bring it under one set of mapping, validation and testing practices as a unified data team, with reconciliation proving each change matches before it ships.

  • Scale it to about 40× today’s data footprint this year: Find the efficiency in our pipeline processes and revisit the decisions underneath them, including OLAP versus OLTP storage and event streaming versus nightly jobs, so the platform absorbs the volume without a proportional increase in cost or run time.

  • Make mapping and identity explainable: Member IDs, eligibility and claims associate to the right person through logged, reviewable decisions, so wrong associations are caught by us and never reach a member.

  • Build the tooling around the pipes: Ship the observability and automation the team works in, including TypeScript web applications and AI where it helps, such as softer matching of names and plans that no exact rule catches.

  • Run it as production software: Alerts on failure and on data-quality regressions, idempotent re-runs, routine backfills, and secure handling of the SSNs and health information in every file we touch.

What You’ll Bring

We understand that your experience is more than just a list of requirements, so we encourage you to apply even if you don't meet all of the following bullet points.

  • Hands-on ownership of a production data pipeline other teams depended on. Building it, inheriting it and keeping it running and evolving, or owning one stage of it (ingestion, or normalization, mapping and validation) all count. Messy external sources and downstream users who noticed when data was wrong are what make it comparable.

  • Python and SQL depth you can exercise without an AI assistant and use to steer one: joins, indexes, query plans, why a query is slow and whether a fix is real.

  • An orchestrator in production (Dagster preferred; Airflow or Prefect are comparable) and Postgres or a comparable relational database on a cloud.

  • Proof habits: reconciliation, data-quality gates, tests you would trust at 3 a.m.

  • AI-assisted development as a daily workflow: you know what to delegate and when the agent is wrong.

  • Care with sensitive data, and clear communication with people who do not write code.

Not required: Dagster specifically, Spark-scale data, a degree, or a years-of-experience number. We believe the right technology for a use case is mappable across several choices; comparable work matters and the stack you did it on does not. Hands-on dbt is a strong plus, not a gate.

Our Stack

Dagster, dbt and Python on Postgres and AWS; TypeScript web applications for observability and automation; Airbyte, Parabola and Salesforce around them; Claude Code, Codex and Cursor for development. These are the tools we prefer today, and we are open to others where a use case calls for it.

What You’ll Get

  • Real ownership on a concentrated data team, where your decisions on mapping, storage and orchestration set the practice as we scale roughly 40× this year.

  • Work that visibly unlocks Nava’s AI features, member experience and broker conversations, with clients who feel the difference.

  • A direct line to the engineers and leaders making product and platform decisions.

  • A remote-first company with a mission to fix healthcare and the tooling to do the best work of your career.

How We Interview

A 15-minute introduction with our recruiter, then an approximately 15-minute AI conversation about a production pipeline you ran and a time you dealt with bad inbound data (voice or typing, camera off). Then two interview days booked together: a 60-minute hands-on session in a small real data pipeline on Day 1, using the AI coding tools you actually use; a 60-minute design conversation and a 60-minute accomplishment discussion on Day 2. Day 2 depends on the Day 1 outcome. We send preparation notes ahead of each step. Same questions for every candidate; we score evidence of comparable work, not polish.

Working at Nava

As a remote-first company, Nava is committed to building a dynamic and inclusive culture where you have the autonomy to thrive. You’ll be supported by cutting-edge technology, a collaborative team, and a shared mission to revolutionize healthcare.

Candidates from all backgrounds are encouraged to apply. We believe that solving America's healthcare problem requires leveraging America's greatest strength: our diversity. Healthcare affects everyone – and a team that includes people from all backgrounds and walks of life will be more effective at driving change than a homogeneous one. We are excited to build that kind of team at Nava.

HQ

Nava Benefits New York, New York, USA Office

228 Park Ave S, New York, New York, United States

Similar Jobs

5 Days Ago
In-Office or Remote
New York City, NY, USA
124K-207K Annually
Senior level
124K-207K Annually
Senior level
Artificial Intelligence • Healthtech • Machine Learning • Natural Language Processing • Biotech • Pharmaceutical
Build and operate production data pipelines supporting analytics, AI, and agentic workflows. Responsibilities include implementing canonical data models, maintaining Databricks or Snowflake platforms, monitoring reliability, responding to incidents, validating healthcare data mappings, improving performance and cost, and documenting architecture. Requires strong SQL and Python skills, cloud data platform experience, ETL/ELT orchestration expertise, and healthcare or pharmaceutical data experience.
Top Skills: Ai/Ml WorkflowsDatabricksEltETLHedisOmopPythonSnowflakeSQL
6 Days Ago
Remote or Hybrid
USA
125K-159K Annually
Mid level
125K-159K Annually
Mid level
AdTech • Automotive • Big Data • Consumer Web
Administer and enhance Edmunds’ Databricks data platform and AWS infrastructure. Build and maintain ETL pipelines, infrastructure-as-code tooling using Terraform or CDK, and operational dashboards, alerts, and reports. Collaborate with business, engineering, analytics, security, and infrastructure teams to support data platform users and AI solutions. Evaluate new technologies, troubleshoot platform issues, and improve operational, cost, and security visibility.
Top Skills: SparkAWSAws CdkDatabricksInfrastructure As Code (Iac)PythonScalaSQLTerraform
7 Days Ago
Remote or Hybrid
USA
120K-180K Annually
Senior level
120K-180K Annually
Senior level
Cloud • Computer Vision • Information Technology • Sales • Security • Cybersecurity
Architects, deploys, and operates CrowdStrike’s data platform infrastructure and services. Responsibilities include administering Airflow and Superset, managing Terraform infrastructure, building CI/CD pipelines, implementing security and access controls, developing DBT data models, cataloging assets in OpenMetadata, writing Python automation, and mentoring engineers. The role requires cloud, container orchestration, SQL, data modeling, compliance, and AI technology expertise, along with eligibility for CJIS clearance.
Top Skills: AIApache AirflowApache SupersetAWSAzureCi/CdDbtDockerGCPIamKubernetesMachine Learning PipelinesOciOpenmetadataPythonSQLTerraform

What you need to know about the NYC Tech Scene

As the undisputed financial capital of the world, New York City is an epicenter of startup funding activity. The city has a thriving fintech scene and is a major player in verticals ranging from AI to biotech, cybersecurity and digital media. It also has universities like NYU, Columbia and Cornell Tech attracting students and researchers from across the globe, providing the ecosystem with a constant influx of world-class talent. And its East Coast location and three international airports make it a perfect spot for European companies establishing a foothold in the United States.

Key Facts About NYC Tech

  • Number of Tech Workers: 549,200; 6% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Capgemini, Bloomberg, IBM, Spotify
  • Key Industries: Artificial intelligence, Fintech
  • Funding Landscape: $25.5 billion in venture capital funding in 2024 (Pitchbook)
  • Notable Investors: Greycroft, Thrive Capital, Union Square Ventures, FirstMark Capital, Tiger Global Management, Tribeca Venture Partners, Insight Partners, Two Sigma Ventures
  • Research Centers and Universities: Columbia University, New York University, Fordham University, CUNY, AI Now Institute, Flatiron Institute, C.N. Yang Institute for Theoretical Physics, NASA Space Radiation Laboratory

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account