Catapult Logo

Catapult

Senior ML/Data Engineer

Reposted 10 Days Ago
In-Office
New York, NY, USA
158K-259K Annually
Senior level
In-Office
New York, NY, USA
158K-259K Annually
Senior level
Own and build the data foundation for an AI sports-performance platform: design data architecture, real-time streaming ingestion, time-series storage, feature store and millisecond feature serving, sport knowledge graph and evaluation/calibration frameworks, ensure tenant-level isolation and provenance, and work closely with domain scientists and AI engineers to deliver trustworthy, production-grade ML infrastructure.
The summary above was generated by AI

Catapult is building the future of sports performance technology, with a mission to Unleash the Potential of every athlete and team on earth.

Since 2006, our solutions have helped more than 5,000 teams around the world make better decisions about athlete health, readiness, performance, and game-day strategy. Our technology is used across the NFL, NBA, NHL, MLS, EPL, AFL, NRL, NCAA, and many other elite sporting organisations.


We are now building the AI layer that brings together the depth of data Catapult has collected over two decades. Our goal is to become an intelligence partner for coaches, athletes, and performance staff, connecting data from sensors, video, sport science, and historical performance to surface insights that practitioners can trust.

We are looking for a Senior ML / Data Engineer to build the data infrastructure that powers this next generation of performance intelligence.

This is a senior production engineering role. You will own significant parts of the architecture that ingest, store, transform, serve, and evaluate athlete performance data. The systems you build will support real-time and machine learning use cases across a global, multi-tenant platform.

You will work closely with data scientists, ML engineers, software engineers, and sports scientists to turn complex performance requirements into reliable, scalable production infrastructure.

This role is suited to an engineer who has spent several years building and operating production systems and is comfortable taking ownership of architecture and technical decisions.


What You'll Do:

  • Design and build production data infrastructure for high-volume athlete performance and sensor data.
  • Build and operate real-time and near-real-time ingestion systems for streaming data.
  • Design storage and data architectures for high-volume time-series data and longitudinal athlete records.
  • Build infrastructure that makes production features and derived metrics available to machine learning systems and AI agents with low latency.
  • Design and implement graph data models and schemas representing relationships between athletes, training loads, injuries, performance, and outcomes.
  • Build data and ML evaluation infrastructure that helps measure model reliability, calibration, and performance across real-world cases.
  • Design systems that maintain strong tenant-level data isolation across clubs and customers.
  • Establish appropriate data provenance, lineage, auditability, and observability across the platform.
  • Work with ML and AI engineers to provide reliable data foundations for model training, inference, and evaluation.
  • Work with sport scientists and domain experts to translate complex requirements into durable production systems.
  • Make pragmatic technology and architecture decisions as the platform evolves.

What You'll Need:

  • 5+ years of full-time professional software or data engineering experience, excluding internships, university placements, coursework, and academic projects.
  • Proven experience designing, building, and operating production data infrastructure at scale.
  • Strong experience working with time-series data or time-series databases, such as InfluxDB, TimescaleDB, Prometheus, ClickHouse, or equivalent technologies.
  • Significant experience with real-time or streaming data ingestion, using technologies such as Kafka, Kinesis, Flink, Spark Streaming, Pulsar, or equivalent.
  • Experience designing graph data models or graph database schemas, not simply querying or consuming an existing graph database.
  • Experience designing or operating multi-tenant systems with tenant-level data isolation.
  • Strong Python and SQL skills. Professional experience with Go is highly desirable.
  • Experience working with production systems where reliability, scalability, observability, and data correctness matter.
  • Ability to take ownership of ambiguous technical problems and turn them into practical production architectures.
  • Experience working directly with data scientists, ML engineers, or other technical domain specialists.
  • Experience building probabilistic evaluation, model calibration, or model monitoring infrastructure.
  • Experience with causal inference, counterfactual modelling, or simulation.
  • Experience working with wearable sensors, IoT data, biomechanics, sports technology, or other high-frequency telemetry.
  • Experience building knowledge graphs or domain-specific ontologies.
  • Experience with LLM or AI evaluation frameworks and an understanding of their limitations.
  • Experience with AWS, including ECS, EC2, Lambda, SNS, SQS, or related services.
  • Experience with GraphQL, REST, gRPC, Postgres, MongoDB, or similar technologies.
  • What We Mean by Senior

This role requires demonstrated professional ownership of production systems.

We are not looking for someone whose primary exposure to these technologies comes from internships, university projects, coursework, or short-term placements.

You do not need experience with every technology listed above. We care more about the depth of your production experience, your ability to design systems, and your track record of taking ownership of complex engineering problems.

For example, strong experience designing and operating Kafka-based streaming infrastructure is more valuable to us than having used five different streaming technologies at a superficial level.

Similarly, we are looking for engineers who have designed graph schemas, not simply listed Neo4j on their CV.


The platform requires capabilities including:

  • Real-time streaming ingestion
  • High-volume time-series storage
  • Data lake and analytical infrastructure
  • Low-latency feature serving
  • Graph databases and domain-specific ontologies
  • ML evaluation and calibration
  • Causal and simulation modelling
  • Data provenance and audit logging
  • Strong tenant-level isolation
  • Production observability and reliability
  • No individual vendor or technology is locked in. We value engineers who understand the underlying architectural trade-offs and can choose the right technology for the problem.

Why Catapult?

Catapult has spent more than twenty years collecting ground-truth athlete data from hardware on the body and on the field, across more than 40 sports and 100 countries.

That data represents a significant opportunity to build new forms of performance intelligence. The challenge is turning that data into systems that are reliable, explainable, and useful at the point where coaches and performance staff need to make decisions.

You will have the opportunity to work on a technically challenging combination of real-time data, machine learning, time-series infrastructure, graph data, and AI evaluation, with direct impact on products used by elite sporting organisations around the world.

Compensation & Benefits

The target Total Compensation range for this position is $157,945 to $259,480 per year.

This range is inclusive of base salary and a target incentive plan, which may include equity, commission, or other bonus structures.
Your specific compensation will be determined by factors including geographic location, relevant experience, and job-related skills.
Catapult also offers paid leave and recognised company holidays, together with a comprehensive benefits package including Health, Dental, Vision, and a 401(k) retirement plan with company match.

Whether you are passionate about sport or simply excited by difficult engineering problems, you will have the opportunity to build technology used by some of the world's most successful teams and athletes.

Catapult is an equal opportunity employer. We value diverse perspectives and encourage people from a wide range of backgrounds to apply.
If you have strong production engineering experience but do not meet every preferred requirement, we would still like to hear from you. We are more interested in depth of experience, technical judgement, and your ability to build reliable systems than in a perfect match against every technology listed.

All offers of employment are subject to Catapult's positive prehire check. To find out more, please contact the Talent Partner for this role.

Similar Jobs

15 Days Ago
In-Office
139K-232K Annually
Senior level
139K-232K Annually
Senior level
Artificial Intelligence • Healthtech • Machine Learning • Natural Language Processing • Biotech • Pharmaceutical
Lead development of an AI-ready research data ecosystem for vaccine R&D: design integrated data architectures, automate ingestion and transformation pipelines, define semantic/metadata frameworks, build knowledge graphs and RAG systems, and develop multimodal predictive and translational ML models. Provide technical leadership, mentor teams, and collaborate across scientists, bioinformatics, and enterprise data organizations to accelerate AI-driven vaccine discovery and decision-making.
Top Skills: Agentic AiCloudFoundation ModelsGenerative AiHigh-Performance ComputingKnowledge GraphsPythonPyTorchRetrieval-Augmented Generation (Rag)TensorFlow
16 Days Ago
In-Office
New York, NY, USA
Senior level
Senior level
Fintech • Financial Services
Lead architecture and delivery of frontier AI solutions: define end-to-end technical designs, establish reusable AI engineering patterns, guide platform integration and production readiness, advise cross-functional stakeholders, and ensure secure, scalable, observable AI products that are commercially viable.
Top Skills: Agent ProtocolsAIAPIsCi/CdCloud InfrastructureDeep LearningFeature PipelinesKnowledge GraphsMachine LearningMemory-Based AgentsMicroservicesMlopsModel InferenceModel Lifecycle ToolingModel ServingModel TrainingObservabilityOrchestrationProduction Ai ObservabilityRetrieval SystemsSecure DeploymentVector Databases
14 Days Ago
In-Office or Remote
United States
42K-148K Annually
Senior level
42K-148K Annually
Senior level
Agency • Information Technology
Design, develop, fine-tune, and deploy generative AI models using deep learning and transformer architectures. Collaborate cross-functionally, troubleshoot model issues, optimize performance, document work, and communicate technical concepts to non-technical stakeholders.
Top Skills: Bert)Deep LearningGansNlpPrompt EngineeringPythonPyTorchTensorFlowTransformers (GptVaes

What you need to know about the NYC Tech Scene

As the undisputed financial capital of the world, New York City is an epicenter of startup funding activity. The city has a thriving fintech scene and is a major player in verticals ranging from AI to biotech, cybersecurity and digital media. It also has universities like NYU, Columbia and Cornell Tech attracting students and researchers from across the globe, providing the ecosystem with a constant influx of world-class talent. And its East Coast location and three international airports make it a perfect spot for European companies establishing a foothold in the United States.

Key Facts About NYC Tech

  • Number of Tech Workers: 549,200; 6% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Capgemini, Bloomberg, IBM, Spotify
  • Key Industries: Artificial intelligence, Fintech
  • Funding Landscape: $25.5 billion in venture capital funding in 2024 (Pitchbook)
  • Notable Investors: Greycroft, Thrive Capital, Union Square Ventures, FirstMark Capital, Tiger Global Management, Tribeca Venture Partners, Insight Partners, Two Sigma Ventures
  • Research Centers and Universities: Columbia University, New York University, Fordham University, CUNY, AI Now Institute, Flatiron Institute, C.N. Yang Institute for Theoretical Physics, NASA Space Radiation Laboratory

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account