Holly Logo

Holly

Senior Software Engineer, Data & AI

Posted 20 Hours Ago
Be an Early Applicant
In-Office
New York, NY, USA
170K-216K Annually
Senior level
In-Office
New York, NY, USA
170K-216K Annually
Senior level
Build and own Holly’s data platform, including ingestion and normalization pipelines for messy government data, canonical datasets, semantic search, retrieval, and inference APIs. Design schemas, data-quality systems, lineage, observability, and automated production workflows. Integrate LLMs, embeddings, and other AI capabilities for extraction, classification, matching, and search. Partner with founders and product engineers to make architecture decisions and deliver reliable, scalable data infrastructure.
The summary above was generated by AI

Holly is the HR platform built for city and county government. We help local governments modernize how they hire, classify, and manage their workforce — work that directly shapes public services for millions of people. Our platform is live across 12 states with 60+ jurisdictions representing over 10% of the US population, including major counties like Santa Clara and Contra Costa in the Bay Area, Orange County in LA, and Snohomish County in Washington State. Demand is outpacing our team, and we're building to meet it.

Holly is a seed-stage team of 20+ operators with deep roots in public service, civic tech, and AI. We’ve raised $10M from leading investors in government technology and the future of work and have grown over 100X in the last year. We center public servants and impact in every decision, and we bring that same care to how we work here at Holly.


About the Role

We're hiring a Senior Software Engineer, Data & AI to build Holly's Data Platform. Local governments publish enormous amounts of public information salary schedules, job classifications, MOUs (Labor Union Agreements), budgets but it's scattered across thousands of websites and buried in messy formats: scanned PDFs, inconsistent HTML, spreadsheets, and everything in between. Turning that chaos into clean, searchable, trustworthy data is one of our biggest product challenges and deepest moats.

You'll design and build the Data Platform: ingestion and processing pipelines, canonical datasets, semantic search and retrieval, and inference APIs that product engineers can build on. AI is part of the infrastructure, not a layer added later. You'll use models where they improve extraction, classification, normalization, matching, and search, then evaluate and operate those systems in production.

This is a hands-on, high-ownership engineering role. You'll own major data systems, make architecture calls, and raise the standards for how Holly collects, models, and trusts its data. You'll work directly with the founders and partner closely with product engineers to make sure they have clean, reliable data to build on.

If you've worked with large, high-volume data and love the challenge of taming messy real-world inputs into something people can rely on, we'd love to talk.

What You'll Do

You'll own major parts of the Data Platform end to end, make high-leverage technical decisions, and build the data and AI infrastructure the rest of the product depends on.

Own and Build the Data Platform

  • Own the data platform end-to-end from raw public sources to clean, canonical datasets the product consumes
  • Design the architecture, schemas, and standards for how Holly ingests, models, and trusts its data
  • Partner with the founders to scope work, make tradeoffs, and drive delivery on our highest-leverage data initiatives
  • Help set the long-term direction for our data foundation

Build Ingestion & Normalization Pipelines

  • Build systems that collect large volumes of public government data from thousands of local-government sources across the web
  • Turn messy, heterogeneous inputs scanned PDFs, inconsistent HTML, spreadsheets into structured, normalized data (parsing, extraction, OCR, dedupe, entity resolution, schema mapping)
  • Where it adds leverage, incorporate LLM-assisted extraction and embeddings into the pipeline
  • Build for freshness, reliability, and scale so data stays current and trustworthy

Model & Serve Data for the Product

  • Design canonical data models and domain schemas that product engineers build on
  • Expose clean, versioned, well-documented datasets the main app can reliably consume
  • Own data quality, validation, lineage, and observability so downstream teams can trust what they're building on

Build Search & AI Infrastructure

  • Build semantic search and retrieval over large, changing government datasets
  • Design inference APIs for extraction, classification, matching, and other AI-powered data processing
  • Build evals, tracing, and fallbacks so model behavior is measurable and dependable
  • Give product engineers clear interfaces for using data and AI capabilities

Make It Automated & Intelligent

  • Evolve pipelines from manual/one-off toward automated, self-healing, monitored systems
  • Establish data-quality checks, alerting, and standards that keep the platform reliable as it grows
  • Raise the bar on how we collect, validate, and serve data across the company
How We Work

Six principles drive how we build:

  • Work on What Matters, Default to No every yes has a cost, so we save them for what moves the business and spend time on what matters.
  • Question Everything, Be Opinionated titles don't settle arguments, the better case wins. Feel empowered to push back to everyone from the Head of Engineering to one of the founders.
  • Obsess Over Craft quality first, and we don't trade it for a date. If you wouldn't put your name on it, it shouldn'''t end up in the codebase.
  • Own It End to End if you build it, you own it: to production, in tests, and when it breaks. With great power comes great responsibility, with autonomy comes responsibility to make sure you own your work.
  • Ship Small, Ship Often the smallest thing that stands on its own, kept reversible. Small ships compound, are easier to review and easier to fix if there are issues.
  • Automate the Hurt, Not the Itch automate the recurring pain, the Toil aka things you do repeatedly that waste time, not the one-off annoyances or what seems ''"fun''" to automate.
What You'll Have

We'd love to talk if you're a senior engineer who can build data infrastructure and production AI systems, and knows how to turn messy real-world inputs into dependable product capabilities.

The Essentials
  • Have 5+ years building and shipping production software (or equivalent experience)
  • Have a strong track record building and operating production data systems
  • Have worked with large-scale, high-volume data ideally where lots of sources, users, or records make volume and reliability matter
  • Are strong at data modeling and SQL, with experience designing schemas that others build on (Postgres a plus)
  • Have built and owned ETL/ELT pipelines that handle messy, heterogeneous, real-world inputs (scraped data, PDFs, HTML, spreadsheets)
  • Have built production search, retrieval, inference, or AI-assisted data-processing systems
  • Know how to evaluate model quality and operate model-backed systems when outputs are probabilistic
  • Bring a strong data-quality mindset validation, testing, monitoring, lineage, and reliability are core to how you work
  • Take ownership and move fast you work independently, ship often, and thrive in early-stage ambiguity
  • Have a growth mindset you learn quickly, and raise the bar through collaboration and clear standards
  • Are pragmatic about tooling and comfortable working in (or ramping quickly into) a modern TypeScript/Postgres codebase
Bonus Points
  • Open source contributions to or maintainer of a widely used tool.
  • Experience with large-scale web scraping / crawling, document extraction (OCR), or LLM-assisted parsing
  • Experience with embeddings / vector search or supporting ML/AI data workflows
  • Experience with analytical/columnar or warehouse stacks (ClickHouse, BigQuery, Snowflake) and/or streaming pipelines
  • Experience in government, public sector, or civic tech
  • Prior early-stage startup experience

Don't meet every bullet? Apply anyway. If you're strong on most of this and excited about the work, we want to hear from you we'll help you ramp on the rest.

What You'll Get
  • End-to-end ownership. Architect and build data systems that define how the product works.
  • Technical influence. Make the high-leverage calls on data architecture, standards, and how we scale.
  • Commitment to Open Source. We are big believers in supporting open source, and provide a monthly day of Open Source where you can work on your favorite tool. In addition to internal hackathons and other projects
  • Direct access. Work directly with the founders, with autonomy to drive major initiatives end-to-end.
  • High-impact scope. Build the data foundation the entire product depends on, and see its power features customers rely on quickly.
  • Public-service impact. Your work improves how local governments operate, helping millions of Americans access public-service careers.
  • Competitive package. $170k-$216k base, 0.15-0.40% equity (L3), comprehensive health benefits (platinum plan with vision and dental), 401(k), paid parental leave, and a professional development stipend.
Ready to Join Us? A few important notes:

Location: This is an onsite role based out of our New York City HQ, four days a week (typically Mondays Thursdays), with some flexibility depending on the role and the candidate. Candidates must reside in New York or be able to commute to our NYC office. Applicants must be authorized to work in the U.S. without requiring sponsorship.

Work Philosophy: We'''re an early-stage startup serving government clients with hard deadlines. There may be occasional off-hours work around launches or critical issues (rare and typically planned). We value flexibility and trust you to manage your schedule while maintaining a high bar for responsiveness and customer outcomes.

We're excited to build with you

Team Holly 🌆

www.hollygov.com

Holly is committed to building a diverse company and working with the broadest talent pool possible. We encourage applications from all races, religions, national origins, genders, sexual orientations, gender identities, gender expressions, and ages, as well as veterans and individuals with disabilities. If you need a reasonable accommodation during the application or interview process, let us know.

Similar Jobs

9 Days Ago
In-Office
New York, NY, USA
180K-240K Annually
Senior level
180K-240K Annually
Senior level
Legal Tech
Design, build, and operate scalable data pipelines and storage/indexing for an AI-driven legal platform. Implement queue-based asynchronous processing, multi-tenant architectures, observability, migration strategies, and system design documentation. Collaborate across engineering, product, and security to ensure data reliability, performance, and operational readiness.
Top Skills: AirflowAzureDagsterDockerGrafanaKubernetesOauth2OidcOpentelemetryPostgresPrometheusPythonQdrantRabbitMQSAMLSolrTemporalTerraform
9 Minutes Ago
Remote or Hybrid
United States
17-25 Hourly
Junior
17-25 Hourly
Junior
Artificial Intelligence • Automotive • Greentech • Information Technology • Machine Learning • Software • Cybersecurity
Provides routine product-use assistance and technical support for Dealertrack products. Troubleshoots and documents system issues, logs customer cases in CRM, manages multiple tickets, meets service-level targets, follows up with customers, and coordinates with internal departments. Requires flexible availability for business-hour shifts, including Saturdays and overtime, along with strong communication, problem-solving, independence, and teamwork.
Top Skills: CRMDealertrack DmsGenesys CloudSalesforce
20 Minutes Ago
Hybrid
178K-297K Annually
Expert/Leader
178K-297K Annually
Expert/Leader
Artificial Intelligence • Automotive • Greentech • Information Technology • Machine Learning • Software • Cybersecurity
Leads enterprise database strategy, administration, optimization, high availability, disaster recovery, and performance initiatives. Stabilizes on-premises databases and designs phased migrations to AWS. Provides technical authority on SQL Server, relational and NoSQL platforms, data architecture, CI/CD, observability, and resilient cloud-native systems. Collaborates with engineering and architecture teams, mentors database administrators, guides modernization efforts, and communicates technical trade-offs to stakeholders.
Top Skills: Always OnAmazon DynamodbAmazon RdsAWSCi/CdDatabase ClusteringDatabase ReplicationGoldengateLog ShippingMicrosoft Sql ServerMySQLOraclePostgresReplica Sets

What you need to know about the NYC Tech Scene

As the undisputed financial capital of the world, New York City is an epicenter of startup funding activity. The city has a thriving fintech scene and is a major player in verticals ranging from AI to biotech, cybersecurity and digital media. It also has universities like NYU, Columbia and Cornell Tech attracting students and researchers from across the globe, providing the ecosystem with a constant influx of world-class talent. And its East Coast location and three international airports make it a perfect spot for European companies establishing a foothold in the United States.

Key Facts About NYC Tech

  • Number of Tech Workers: 549,200; 6% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Capgemini, Bloomberg, IBM, Spotify
  • Key Industries: Artificial intelligence, Fintech
  • Funding Landscape: $25.5 billion in venture capital funding in 2024 (Pitchbook)
  • Notable Investors: Greycroft, Thrive Capital, Union Square Ventures, FirstMark Capital, Tiger Global Management, Tribeca Venture Partners, Insight Partners, Two Sigma Ventures
  • Research Centers and Universities: Columbia University, New York University, Fordham University, CUNY, AI Now Institute, Flatiron Institute, C.N. Yang Institute for Theoretical Physics, NASA Space Radiation Laboratory

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account