Develop novel machine learning methods and architectures for generative conversational speech-to-speech models. Research speech synthesis and recognition, improve model quality and realism, collaborate on data and infrastructure, and help scale proven research into production pipelines and Spotify products.
he Personalization team makes deciding what to play next easier and more enjoyable for every listener. From Blend to Discover Weekly, we're behind some of Spotify's most-loved features. We built them by understanding the world of music and podcasts better than anyone else. Join us and you'll keep millions of users listening by making great recommendations to each and every one of them.
Within Personalization, the Speak Team owns the development of Spotify's state-of-the-art speech models, contributing to speech recognition, speech synthesis, and speech-to-speech models. We craft voice models that match human-level emotional expressiveness, so we can deeply engage our listeners and support creators at scale. Our groundbreaking work on speech synthesis relies on state-of-the-art deep learning methods and evaluation techniques, highly efficient data processing and model serving, and capturing audio of outstanding quality from our voice talent pool.
We're looking for a senior applied research scientist with experience in developing novel ML techniques and architectures and with a strong interest in working across a full production pipeline to produce state-of-the-art generative conversational speech-to-speech models. You'll collaborate with our engineering teams to help develop our production pipelines, explore new ideas and methods to improve quality, understanding and realism, as well as push the frontiers of what is possible with our speech technology.
What You'll Do
- Develop and experiment with new methods for speech synthesis and speech recognition, along with end-to-end approaches, building on the latest research and ideas.
- Work towards the expansion of our speech use-cases targeting different markets and products.
- Be part of a highly motivated research team dedicated to building and creating models at scale to power the Spotify platform.
- Champion best practices for research and development, sharing your knowledge and experience with other researchers within Speak.
- Collaborate with our engineering and data teams on ideas requiring new infrastructure or new high-quality data, as well as to help improve our speech recognition and speech synthesis pipelines, and help turn proven ideas into scalable products.
Who You Are
- You have a strong background in ML (PhD degree on top of professional experience), and
- experience in working with any of the following: transformers, GANs, diffusion models, flow matching, VAEs, audio codecs.
- You have experience in developing generative models for speech synthesis, speech recognition, audio/music, natural language processing, or computer vision.
- You have strong experience with Python, particularly PyTorch.
- You have strong communication skills and the ability to explain technical ideas with clarity to technical and non-technical people alike.
- You have experience in an academic or professional setting conducting high-quality research.
Where You'll Be
- This role is based in New York City.
- We offer you the flexibility to work where you work best! There will be some in person meetings, but still allows for flexibility to work from home
The United States base range for this position is $169,157 - $241,653 plus equity. The benefits available for this position include health insurance, six month paid parental leave, 401(k) retirement plan, a monthly meal allowance, 23 paid days off, 13 paid flexible holidays. These ranges may be modified in the future.
Spotify is an equal opportunity employer. You are welcome at Spotify for who you are, no matter where you come from, what you look like, or what’s playing in your headphones. Our platform is for everyone, and so is our workplace. The more voices we have represented and amplified in our business, the more we will all thrive, contribute, and be forward-thinking! So bring us your personal experience, your perspectives, and your background. It’s in our differences that we will find the power to keep revolutionizing the way the world listens.
At Spotify, we are passionate about inclusivity and making sure our entire recruitment process is accessible to everyone. We have ways to request reasonable accommodations during the interview process and help assist in what you need. If you need accommodations at any stage of the application or interview process, please let us know - we’re here to support you in any way we can.
Spotify New York, New York, USA Office
4 World Trade Center, New York, NY, United States, 10007
Similar Jobs
Edtech • Social Impact
Own and improve the core learner platform to increase learning rate, retention, and engagement at scale. Prioritize an AI-first roadmap, ship AI-assisted features (AI Tutor, AI Grader), define outcome metrics, run experiments, and collaborate with Curriculum, Learning, and Engineering to deliver high-quality, engaging learner experiences.
Top Skills:
AIAmplitudeClaude CodeCodexCursorFigmaFramerHexMixpanelSQL
Artificial Intelligence • Big Data • Cloud • Information Technology • Software • Big Data Analytics • Automation
Lead enterprise data architecture across warehouses, data lakes, downstream applications, and AI workloads. Design data models, Medallion Architecture, ELT standards, semantic layers, APIs, access controls, and reusable datasets. Partner with business and BI teams to drive adoption, conduct architecture reviews, establish platform standards, and mentor data and analytics engineers. The role requires deep Snowflake, dbt, data modeling, pipeline, and semantic-layer expertise, with experience in AI data patterns and cloud infrastructure preferred.
Top Skills:
AWSAzureDatabricksDbtDelta LakeGCPGraphQLLlmsPower BIRagRestSnowflakeSnowflake Unity CatalogTableau
33 Minutes Ago
Artificial Intelligence • Healthtech • Logistics • Social Impact • Software • Telehealth
Engages patients through high-volume inbound and outbound calls, educates them about ordered at-home healthcare services, schedules appointments, answers questions, and escalates concerns. The role requires approximately 150–200 outbound calls and 80–100 patient conversations daily, collaboration with healthcare professionals, and proficiency with EHR and healthcare software. Zendesk and Five9 experience are preferred.
Top Skills:
Auto-Dialer SystemsElectronic Health Records (Ehr)Five9Zendesk
What you need to know about the NYC Tech Scene
As the undisputed financial capital of the world, New York City is an epicenter of startup funding activity. The city has a thriving fintech scene and is a major player in verticals ranging from AI to biotech, cybersecurity and digital media. It also has universities like NYU, Columbia and Cornell Tech attracting students and researchers from across the globe, providing the ecosystem with a constant influx of world-class talent. And its East Coast location and three international airports make it a perfect spot for European companies establishing a foothold in the United States.
Key Facts About NYC Tech
- Number of Tech Workers: 549,200; 6% of overall workforce (2024 CompTIA survey)
- Major Tech Employers: Capgemini, Bloomberg, IBM, Spotify
- Key Industries: Artificial intelligence, Fintech
- Funding Landscape: $25.5 billion in venture capital funding in 2024 (Pitchbook)
- Notable Investors: Greycroft, Thrive Capital, Union Square Ventures, FirstMark Capital, Tiger Global Management, Tribeca Venture Partners, Insight Partners, Two Sigma Ventures
- Research Centers and Universities: Columbia University, New York University, Fordham University, CUNY, AI Now Institute, Flatiron Institute, C.N. Yang Institute for Theoretical Physics, NASA Space Radiation Laboratory
.png)


