Avaak Logo

Avaak

Data Scientist

Reposted 23 Days Ago
In-Office
New York City, NY, USA
Mid level
In-Office
New York City, NY, USA
Mid level
Perform end-to-end data science for underwriting, auto-quoting, and sales analytics: detect emerging risks, quantify data blind spots, build multi-signal models and confidence scoring, measure algorithmic impact on sales, and maintain ingestion and pipeline data quality and monitoring.
The summary above was generated by AI

Most of what makes American healthcare expensive isn’t medical care. It’s the machinery wrapped around it: middlemen taking a cut, fraud nobody stops, and billing systems designed to fight over payment instead of deliver care. The result is higher premiums, denied claims, surprise bills, and a system patients increasingly experience as adversarial.

Arlo is rebuilding health insurance for small businesses from first principles: making sure as much of every premium dollar as possible goes to care instead of getting absorbed by the system around it. We do that by identifying fraud earlier, steering members toward higher-quality and lower-cost care, automating operational overhead, and eliminating vendors whose business exists mostly to take a cut.

AI is the foundation that makes this work. We use it across underwriting, operations, clinical programs, and member experience to build an insurer that becomes more efficient as the technology improves.

We’re already operating at meaningful scale: profitable, hundreds of millions in premiums, tens of thousands of members covered, and growing quickly through brokers, employers, and partners. Backed by Upfront Ventures, 8VC, and General Catalyst, with a team from Palantir, YC companies, and longtime healthcare operators.

About the role

Arlo quotes small businesses using AI-powered underwriting that predicts individual risk to determine group rates. But accurate pricing requires more than historical data alone — at the time of quoting, additional signals like prior rates, aggregate claims history, and self-reported health information can meaningfully sharpen our view of a group's risk profile.

This role sits at the center of four connected problems: learning from post-policy data to understand where our risk blindness lies, building automated quoting systems that act on softer signals, understanding how algorithmic changes ripple through to sales outcomes, and maintaining the integrity of the data pipelines that underpin all of it.


What you'll work on

Post-quoting intelligence
  • Emerging risk detection. Analyze early-period claims data to surface patterns that indicate risk was underestimated at quoting — and quantify the gap.

  • Blindness measurement. Build frameworks to systematically identify where the external claims database is incomplete or lagged, and estimate the magnitude of those gaps.

  • Rate determination. Develop calibration factors that translate soft signals — based on competitive information, econometrics and emerging events — into rate adjustments layered on top of core model output.

  • Feedback loop infrastructure. Create pipelines that carry post-quoting learnings back into upstream models so calibration improves continuously.

Auto-quoting automation
  • Multi-signal fusion. Design models that ingest heterogeneous inputs and synthesize them into an enriched risk view that reduces reliance on manual review.

  • Confidence scoring. Build the logic that decides when a quote can be issued automatically versus routed to a human, minimizing queue volume without increasing adverse selection.

  • Threshold and rule design. Collaborate with underwriting to set and validate auto-issuance decision thresholds, and monitor their performance over time.

Sales performance analytics
  • Algorithm impact attribution. When we adjust rating algorithms, tighten auto-quoting rules, or change data quality requirements, measure the downstream effect on quote volume, win rates, and sales conversion — and distinguish model-driven changes from market-driven ones.

  • Sales funnel diagnostics. Identify where quote kickouts, rate changes, or data submission friction are creating drop-off in the sales pipeline, and quantify the cost of each leakage point.

  • Data quality incentive analysis. Understand whether brokers and groups that provide richer data at submission (more complete census, aggregate claims, prior rates) achieve better pricing outcomes — and help us make that case externally.

Data ingestion and pipeline integrity
  • Ingest delay profiling. Map the latency characteristics of our live policy data — claims, enrollment, eligibility — and identify which delays are structural versus operational, and where they introduce systematic bias into our models.

  • Consistency monitoring. Build checks and alerting for data inconsistencies across our ingestion layer: duplicate records, mismatched member IDs, enrollment timing gaps, and reporting lags from carriers or TPAs.

  • Upstream data partnerships. Work with engineering and data teams to document known data quality issues, prioritize fixes, and maintain a clear picture of what the data can and cannot reliably support at any given point.

What we're looking for

  • 3–5 years in a data science or quantitative analyst role

  • Proficiency in Python (scikit-learn, pandas, statsmodels) and SQL

  • Experience building and validating predictive models end-to-end

  • Comfort working with messy, inconsistent real-world datasets

  • Experience designing or auditing data pipelines for quality and consistency

  • Ability to communicate model behavior and business impact clearly to non-technical stakeholders

Nice to have

  • Background in insurance, actuarial data, or healthcare claims

  • Experience with GLMs, survival analysis, or credibility theory for pricing

  • Familiarity with group health underwriting or broker distribution models

  • Experience building confidence-based routing or triage systems

  • Exposure to sales funnel analytics or conversion attribution

  • Familiarity with MLflow, dbt, or similar tooling


How you'll work

  • You'll be an individual contributor embedded with underwriting, pricing, and sales teams — with direct access to the people who use your outputs daily.

  • You'll own projects from data exploration through to production deployment and monitoring, without a separate ML engineering handoff layer.

  • Work spans Python-based modeling and SQL-driven analysis equally; both matter.

  • Your models directly influence pricing decisions and sales outcomes, so accuracy, explainability, and sound uncertainty quantification are all first-class concerns.

  • You'll work closely with engineering on data ingestion questions — this role requires comfort sitting at the boundary between data science and data engineering.

Why Join Arlo:

  • High ownership: You’ll get real responsibility from day one—our high-trust team empowers you to run with big problems and shape core parts of the company.

  • Join an important mission: Your work directly influences how people access care and improves lives at scale.

  • Growth & expansion: We’re moving fast, and as we grow, your scope will grow with us—new challenges, bigger opportunities, and rapid career velocity.

  • Apply AI to a problem that matters: Instead of optimizing ads or cutting labor costs, you’ll use AI to fundamentally reimagine how people get healthcare.

  • High pace, high collaboration: We operate with velocity, first-principles thinking, and a team that works closely, openly, and with ambition.


Exact compensation inclusive of salary and any bonuses is determined based on a number of factors including experience and skill level, location, and qualifications which are assessed during the interview process.
Arlo is an equal opportunity employer. We do not discriminate based on age, race, color, creed or religion, national origin, sexual orientation, gender identity or expression, military status, sex, disability, predisposing genetic characteristics, marital status, familial status, status as a victim of domestic violence, or arrest or conviction record, as defined under New York State law.
🔒 Your safety matters to us. If you're selected to move forward in our hiring process, you'll hear directly from a member of our Recruiting team via an @joinarlo.com email address. We will never ask for personal or financial information outside of our formal onboarding process. When in doubt, please reach out to us to verify at: [email protected].

Similar Jobs

3 Days Ago
Hybrid
New York, NY, USA
Expert/Leader
Expert/Leader
Financial Services
Lead Card Data & Analytics strategy and a multi-layered analytics organization to deliver AI/ML solutions (including generative and agentic AI). Set analytical direction, partner with cross-functional teams to productionize models, drive experimentation and causal measurement, build competitive intelligence, and develop talent and operating practices to inform product, pricing, and customer experience decisions.
Top Skills: Agentic AiCausal InferenceDatabricksExperimentation A/B TestingGenerative AiLarge Language ModelsMlopsPythonRRetrieval-Augmented GenerationSnowflake
3 Days Ago
Remote or Hybrid
USA
120K-180K Annually
Senior level
120K-180K Annually
Senior level
Cloud • Computer Vision • Information Technology • Sales • Security • Cybersecurity
Lead research and engineering of production-grade LLM/GenAI systems for cybersecurity: design models and architectures, drive data labeling and evaluation, collaborate with product and engineering to deploy scalable solutions, mentor data scientists, set technical strategy, and advance AI safety, red-teaming, and applied ML innovations for detection and operational use.
Top Skills: Cloud TechnologiesDeep Learning FrameworksGenaiGpu TechnologiesLlmsPython
4 Days Ago
Remote or Hybrid
USA
149K-187K Annually
Senior level
149K-187K Annually
Senior level
Cloud • Fintech • Information Technology • Machine Learning • Software • App development • Generative AI
Lead advanced analytics and ML efforts to extract insights from large-scale historical data, design features and data representations for models and LLM/agentic systems, partner with product and engineering to build data pipelines and prototypes, evaluate AI workflows, and drive innovation through experiments and empirical analysis to inform product strategy.
Top Skills: Agentic Ai SystemsAws-Native Data And Ml StackBigQueryData PipelinesEmbeddingsFeature EngineeringGCPGraph DatabasesLarge Language Models (Llms)NoSQLRetrieval StrategiesSnowflakeSQLVertex Ai

What you need to know about the NYC Tech Scene

As the undisputed financial capital of the world, New York City is an epicenter of startup funding activity. The city has a thriving fintech scene and is a major player in verticals ranging from AI to biotech, cybersecurity and digital media. It also has universities like NYU, Columbia and Cornell Tech attracting students and researchers from across the globe, providing the ecosystem with a constant influx of world-class talent. And its East Coast location and three international airports make it a perfect spot for European companies establishing a foothold in the United States.

Key Facts About NYC Tech

  • Number of Tech Workers: 549,200; 6% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Capgemini, Bloomberg, IBM, Spotify
  • Key Industries: Artificial intelligence, Fintech
  • Funding Landscape: $25.5 billion in venture capital funding in 2024 (Pitchbook)
  • Notable Investors: Greycroft, Thrive Capital, Union Square Ventures, FirstMark Capital, Tiger Global Management, Tribeca Venture Partners, Insight Partners, Two Sigma Ventures
  • Research Centers and Universities: Columbia University, New York University, Fordham University, CUNY, AI Now Institute, Flatiron Institute, C.N. Yang Institute for Theoretical Physics, NASA Space Radiation Laboratory

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account