Sunset Logo

Sunset

Backend Engineer — Data Pipeline

Posted 9 Days Ago
In-Office
New York, NY, USA
Mid level
In-Office
New York, NY, USA
Mid level
Design and operate backend systems for high-volume, multi-stage data pipelines that de-identify sensitive enterprise data. Own versioned contracts, lineage, state transitions, retries, replay, reconciliation, verification, migrations, and safe recovery. Diagnose production behavior across services and storage, improve correctness and throughput, and build auditable release controls. Collaborate with ML, data, platform, security, product, and customer-facing teams while using AI tools with deterministic safeguards and human review.
The summary above was generated by AI
About Sunset

At its core, Sunset was founded to help founders. We started by supporting startups through shutting down, but we have since expanded into unlocking a new revenue stream for all types of businesses.

In 2025, we had a unique insight: the data every company generates each day through collaboration, communication, and building is some of the most valuable training data in the world. Public and synthetic data can only get frontier models so far, so the next generation of model progress depends on real, proprietary data grounded in how actual businesses operate. We are a primary source of it, partnering directly with the frontier AI labs building what comes next.

Why Join Sunset Now
  • We have scaled from $0 to a multi-eight-figure run rate in a matter of months

  • We have raised from top-tier investors, including Floodgate, Afore, Ludlow, and Hustle Fund

  • We are small enough that you will carry outsized responsibility and grow as quickly as the company does

  • You will partner with and build for some of the fastest and most important companies in the world

  • You will help build a massive, category-defining business from the ground floor

The Role

We're hiring a backend engineer to make the pipeline that de-identifies sensitive enterprise data correct, replayable, operable, and safe to change.

This is not a conventional data-platform role where a successful job is enough. A pipeline can finish while silently dropping records, duplicating output, applying stale policy, losing lineage, or producing evidence that cannot establish whether a dataset is safe to release. You will own the backend systems and contracts that make those failure modes visible, preventable, and recoverable.

You will work across asynchronous orchestration, batch workers, queues, object storage, databases, many file formats, model-backed stages, deterministic verification, and human review. The role is backend-focused, but the outcome is a product and delivery promise: the team must know what ran, what changed, what remains uncertain, and what can safely happen next.

What You'll Own
  • Design and ship backend systems for multi-stage, high-volume data processing

  • Define authoritative, versioned contracts for manifests, artifacts, lineage, and state transitions

  • Make retries, checkpoints, partial failures, replay, backfills, migrations, and rollbacks safe and understandable

  • Build independent reconciliation and verification instead of treating job success as proof of correct output

  • Turn escaped and recurring failures into fixtures, regression coverage, release gates, and durable recovery paths

  • Expose trustworthy run state and safe controls to the products people use to investigate and release data

  • Diagnose production behavior across code, queues, stores, artifacts, data formats, and deployed versions

  • Improve correctness, throughput, and operating leverage without weakening privacy, security, or release confidence

  • Use AI deeply in development and in bounded verification systems, with explicit evaluation and independent checks

  • Partner with machine learning, applied science, full-stack product, platform, security, and data engineering teammates

What Success Looks Like
  • A high-risk pipeline boundary has an explicit contract, independent reconciliation, replayable coverage, and safe recovery

  • Missing, duplicated, stale, or incompatible work is detected before it becomes a customer delivery

  • Material pipeline state and release decisions are backed by queryable provenance and audit evidence

  • Recurring reruns, manual interventions, and diagnosis or recovery time decline

  • Adjacent engineers can add stages and checks through supported patterns instead of one-off scripts and implicit storage conventions

  • Pipeline changes can be rolled out, backfilled, quarantined, or reversed without delivery heroics

You Might Thrive Here If
  • You have at least three years of professional software engineering experience, including personal ownership of production backend systems

  • You are strong in asynchronous or distributed systems and can reason precisely about queues, concurrency, state, storage, idempotency, partial failure, and recovery

  • You have worked on systems where output could be materially wrong even when every service looked healthy

  • You define invariants and use reconciliation, control totals, diffs, replay, goldens, shadow paths, or independent sources to verify correctness

  • You can design versioned data and artifact contracts and migrate them safely in a live system

  • You debug from evidence across system boundaries and turn incidents into durable system improvements

  • You choose technical work based on operator and customer consequences, not architecture in isolation

  • You use modern AI engineering tools fluently, verify their output, and know when model-backed checks need deterministic guardrails and human review

  • You communicate clearly across product, ML, data, platform, security, and customer-facing teams

This Role May Not Be for You If
  • You want to focus primarily on frontend product development or visual craft

  • You treat pipeline success, uptime, latency, or a green dashboard as sufficient evidence that the output is correct

  • You prefer isolated infrastructure work without responsibility for data and delivery consequences

  • You solve partial failure primarily with retries and manual runbooks

  • You do not want AI tools to be part of your daily engineering workflow

Bonus
  • Experience with large-scale batch processing, workflow orchestration, event-driven systems, or data movement

  • Experience with schema evolution, manifests, lineage, CDC, migrations, reindexing, or backfills

  • Experience in payments, ledgers, reconciliation, claims, fraud, identity, search quality, observability, or another domain with delayed or weak ground truth

  • Experience with sensitive or multi-tenant data, least-privilege systems, auditability, quarantine, and fail-closed release paths

  • Experience combining deterministic checks, synthetic fixtures, offline replay, model-based judges, and human review

  • Experience with Python, AWS, Airflow, Batch, SQS, S3, DynamoDB, PostgreSQL, or comparable systems

HQ

Sunset New York, New York, USA Office

Dumbo, New York, NY, United States, 11201

Similar Jobs

56 Minutes Ago
Hybrid
New York, NY, USA
19-38 Hourly
Junior
19-38 Hourly
Junior
eCommerce • Fashion • Retail • Sales • Wearables • Design
Leads store operations alongside the Store Leader, driving sales, profitability, customer service, omnichannel selling, team development, recruiting, coaching, scheduling, visual standards, inventory accuracy, loss prevention, and operational compliance. Analyzes performance metrics, creates action plans, manages customer engagement and livestream shopping, and maintains an inclusive, high-performing retail environment. Requires flexible availability, physical ability to handle stockroom duties, and at least two years of retail management experience.
Top Skills: ExcelMicrosoft OutlookMicrosoft PowerpointMicrosoft WordRetail Management SystemsSocial Media
Yesterday
Hybrid
New York, NY, USA
273K-341K Annually
Expert/Leader
273K-341K Annually
Expert/Leader
Cloud • Information Technology • Security • Software • Cybersecurity
Lead strategic technical relationships for 1–5 high-value enterprise accounts. Develop platform adoption roadmaps across application modernization, developer platforms, and AI; coordinate solution engineers, customer success, and product teams; engage CTOs and engineering leaders; identify expansion opportunities; and communicate customer requirements into product planning. The role requires strong cloud, security, AI, architectural, consulting, and executive communication expertise, with a focus on driving adoption and revenue influence rather than hands-on engineering.
Top Skills: Ai GatewayAWSCloudflare D1Cloudflare Durable ObjectsCloudflare PagesCloudflare QueuesCloudflare R2Cloudflare WorkersGCPMicrosoft 365AzureSecurity ArchitectureVectorizeWorkers AiWorkers Kv
Yesterday
Hybrid
New York, NY, USA
Senior level
Senior level
Financial Services
Develop and scale customer journeys for offers and shopping, translating customer insights and product signals into audience strategies, messaging, campaign concepts, creative briefs, experimentation roadmaps, and optimization recommendations. Partner with analytics, product, creative, marketing technology, and campaign teams to launch initiatives, measure performance, improve activation and redemption, and manage partner campaign workflows and governance.

What you need to know about the NYC Tech Scene

As the undisputed financial capital of the world, New York City is an epicenter of startup funding activity. The city has a thriving fintech scene and is a major player in verticals ranging from AI to biotech, cybersecurity and digital media. It also has universities like NYU, Columbia and Cornell Tech attracting students and researchers from across the globe, providing the ecosystem with a constant influx of world-class talent. And its East Coast location and three international airports make it a perfect spot for European companies establishing a foothold in the United States.

Key Facts About NYC Tech

  • Number of Tech Workers: 549,200; 6% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Capgemini, Bloomberg, IBM, Spotify
  • Key Industries: Artificial intelligence, Fintech
  • Funding Landscape: $25.5 billion in venture capital funding in 2024 (Pitchbook)
  • Notable Investors: Greycroft, Thrive Capital, Union Square Ventures, FirstMark Capital, Tiger Global Management, Tribeca Venture Partners, Insight Partners, Two Sigma Ventures
  • Research Centers and Universities: Columbia University, New York University, Fordham University, CUNY, AI Now Institute, Flatiron Institute, C.N. Yang Institute for Theoretical Physics, NASA Space Radiation Laboratory

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account