As a Data Scientist/Machine Learning Engineer, you'll finetune models, improve data quality, add new signals, and push solutions to production.
Sumble is building a knowledge graph from web data with a first focus on data for go-to-market teams. We use sources like job posts and resume data to identify things like org structure, tech stack, and key projects (e.g., GenAI initiatives, cloud migrations). Our product already has strong product-market fit, early revenue, and happy customers — and now we’re ready to accelerate.
Our long-term vision is to become the primary destination for accessing high-quality web data. Try the product at sumble.com.
Our Team: We are a team of 15, including 10 engineers with experience at companies such as Google, Meta, Stack Overflow, and Kaggle.
What you'll do- Finetuning small language models
- Improving the quality of existing data using scalable approaches. Examples include: making sure URLs are associated the right company, we have the correct HQ address, we have mapped parents-subsidiary using techniques like LLM validation, SERP, and triangulating across sources.
- Adding new signals: this usually involves scrubbing, matching and normalizing new signals and matching to our existing ontology
- Pushing solutions into production environments, which may involve touching data pipelines and/or backend systems
Requirements
- Located within Americas timezones
Our Tech Stack:
- ML/Data: PyTorch, Huggingface, Gemma models, LORA, VLLM, Skypilot, Marimo
- Languages & Frameworks: Python, FastAPI, React, Typescript
- Cloud Platform: Google Cloud Platform (GCP)
- Databases: PostgreSQL, DuckDB
- Infrastructure: Cloud Run
- Product/Design: Figma, Vercel V0
Challenges We Tackle:
- Transforming noisy datasets into high-quality data products
- Running expensive analytics computations efficiently
- Managing the complexity of a growing number of data sources, machine learning models, and large data operations
- Create a great PLG experience with upsell pathways
Benefits
- Medical, dental, and vision (US)
- 401k (US)
- Target 4 weeks PTO
Similar Jobs
Artificial Intelligence • Information Technology • Software
As a founding Data Scientist/Machine Learning Engineer, you'll develop AI/ML models, enhance product capability, and drive impactful user outcomes while working closely with product teams.
Top Skills:
Data ScienceMachine Learning
Fintech • Financial Services
Provide second-line ERM oversight of enterprise fraud risk across products, channels, technologies, and third parties. Develop and maintain fraud risk assessments, map risks to controls, assess control design and effectiveness, track remediation, support analytics and AI/ML-enabled fraud monitoring, prepare executive and board reporting, support regulatory exams and governance, and partner cross-functionally on innovation and vendor risk.
Top Skills:
Ai/MlArcherFraud AnalyticsMetricstreamMicrosoft Office SuiteServicenow Grc
Artificial Intelligence • Cloud • Information Technology • Security • Social Impact • Software
Serve as a technical advisor for assigned accounts, guiding implementations, integrations, and escalations. Collaborate cross-functionally to resolve technical issues, monitor account health, run reviews, identify expansion opportunities, and improve documentation and processes to support customer success.
Top Skills:
Ai ToolsAPIsScimSftpSso
What you need to know about the NYC Tech Scene
As the undisputed financial capital of the world, New York City is an epicenter of startup funding activity. The city has a thriving fintech scene and is a major player in verticals ranging from AI to biotech, cybersecurity and digital media. It also has universities like NYU, Columbia and Cornell Tech attracting students and researchers from across the globe, providing the ecosystem with a constant influx of world-class talent. And its East Coast location and three international airports make it a perfect spot for European companies establishing a foothold in the United States.
Key Facts About NYC Tech
- Number of Tech Workers: 549,200; 6% of overall workforce (2024 CompTIA survey)
- Major Tech Employers: Capgemini, Bloomberg, IBM, Spotify
- Key Industries: Artificial intelligence, Fintech
- Funding Landscape: $25.5 billion in venture capital funding in 2024 (Pitchbook)
- Notable Investors: Greycroft, Thrive Capital, Union Square Ventures, FirstMark Capital, Tiger Global Management, Tribeca Venture Partners, Insight Partners, Two Sigma Ventures
- Research Centers and Universities: Columbia University, New York University, Fordham University, CUNY, AI Now Institute, Flatiron Institute, C.N. Yang Institute for Theoretical Physics, NASA Space Radiation Laboratory



