Stability AI Logo

Stability AI

Generative AI Inference Engineer

Sorry, this job was removed at 06:09 p.m. (EST) on Monday, Mar 09, 2026
Remote
Hiring Remotely in United States
Remote
Hiring Remotely in United States

Similar Jobs

4 Hours Ago
Easy Apply
Remote
USA
Easy Apply
167K-200K Annually
Senior level
167K-200K Annually
Senior level
Big Data • Healthtech • HR Tech • Machine Learning • Software • Telehealth • Big Data Analytics
The Senior Data Engineer will build and maintain data pipelines, create reusable data sets, and ensure data privacy and security, contributing to healthcare innovations.
Top Skills: AirbyteAirflowArgoAWSDbtDuckdbElasticsearchIcebergPostgres/SqlPythonSnowflakeSparkTerraform
14 Hours Ago
In-Office or Remote
Bingen, WA, USA
184K-253K Annually
Senior level
184K-253K Annually
Senior level
Aerospace • Information Technology • Software • Cybersecurity • Design • Defense • Manufacturing
Lead growth initiatives and partnerships in the US Domestic aerospace and defense sectors, focusing on business development and customer engagement.
Top Skills: Microsoft Office SuiteSalesforce
16 Hours Ago
In-Office or Remote
Salt Lake City, UT, USA
67K-106K Annually
Entry level
67K-106K Annually
Entry level
Cloud • Information Technology • Productivity • Security • Software • App development • Automation
The Sales Development Representative will manage outbound and inbound leads, build relationships, and collaborate with account executives to improve sales efforts.
Top Skills: B2BMarketingSaaSSales

Generative AI Inference Engineer

<Remote> 

About the role: 

We are seeking passionate Machine Learning Engineers to join our Inference team, focusing on the creative applications of generative AI models. The ideal candidate will have substantial experience developing and running inference for multi-modal models. A deep understanding of diffusion model architectures and familiarity with workflow tools like ComfyUI are a big plus. You will be expected to leverage and push the boundaries of state-of-the-art inference optimization techniques for multi-modal generative models. This role offers the opportunity to work alongside top researchers and engineers, utilizing cutting-edge high-performance computing resources to make a significant impact in the rapidly evolving field of generative AI.

Responsibilities:  

  • Lead efforts to drive the design, development of customer-facing multi modal ML inference systems.
  • Work with the Platform and Inference teams on building inference systems for the next generation of models, where you will work on areas such as optimization, model tuning and deployment.
  • Partner with leading cloud providers to deliver hosted Stability AI inference solutions.
  • Be a strategic thought partner for leaders across the organization on driving business impact through machine learning
  • Be part of the team to bring new Stability models and pipelines into existence
  • Prototype and productionize inference platform improvements and new features 

Qualifications:

  • 7+ years working on productionizing machine learning systems, including inference pipeline development
  • Expert level knowledge on writing and running python services at scale
  • 5+ years working on python scientific stack, pyTorch and at least one high-performance inference framework (e.g. Triton and TensorRT)
  • Deep understanding of Diffusion Architecture
  • Experience profiling and optimizing deep neural networks on Nvidia GPUs, using profiling tools such as NVIDIA Nsight
  • Experience with python-based image manipulation/encoding/decoding frameworks, such as OpenCV
  • Experience deploying to cloud orchestration systems such as Kubernetes and cloud providers such as AWS, GCP, and Azure
  • Experience with Docker
  • Ability to rapidly prototype solutions and iterate on them with tight product deadlines
  • Strong communication, collaboration, and documentation skills
  • Experience with the open-source ML ecosystem (HuggingFace, W&B, etc.)

Equal Employment Opportunity:

We are an equal opportunity employer and do not discriminate on the basis of race, religion, national origin, gender, sexual orientation, age, veteran status, disability or other legally protected statuses.

What you need to know about the NYC Tech Scene

As the undisputed financial capital of the world, New York City is an epicenter of startup funding activity. The city has a thriving fintech scene and is a major player in verticals ranging from AI to biotech, cybersecurity and digital media. It also has universities like NYU, Columbia and Cornell Tech attracting students and researchers from across the globe, providing the ecosystem with a constant influx of world-class talent. And its East Coast location and three international airports make it a perfect spot for European companies establishing a foothold in the United States.

Key Facts About NYC Tech

  • Number of Tech Workers: 549,200; 6% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Capgemini, Bloomberg, IBM, Spotify
  • Key Industries: Artificial intelligence, Fintech
  • Funding Landscape: $25.5 billion in venture capital funding in 2024 (Pitchbook)
  • Notable Investors: Greycroft, Thrive Capital, Union Square Ventures, FirstMark Capital, Tiger Global Management, Tribeca Venture Partners, Insight Partners, Two Sigma Ventures
  • Research Centers and Universities: Columbia University, New York University, Fordham University, CUNY, AI Now Institute, Flatiron Institute, C.N. Yang Institute for Theoretical Physics, NASA Space Radiation Laboratory

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account