Together AI Logo

Together AI

Technical Compute Qualification Manager

Reposted One Month Ago
Be an Early Applicant
In-Office
New York, NY, USA
200K-250K Annually
Senior level
In-Office
New York, NY, USA
200K-250K Annually
Senior level
Lead and improve the end-to-end qualification process for new compute capacity, running parallel provider evaluations across compute, networking, storage, power, cooling, and operations. Coordinate engineering teams, review specs and test data, perform first-pass analyses, and deliver clear go/no-go recommendations and auditable evaluation records to inform sourcing decisions.
The summary above was generated by AI
About The Role

Together AI is growing its compute footprint, and making sure new capacity meets our technical standards is an important priority for the company. Every new cluster has to clear a technical bar before it carries customer workloads, and this role owns that bar. As Technical Compute Qualification Manager, you will run the process that screens and qualifies prospective compute providers, taking each prospective  deployment through a structured evaluation across compute, networking, storage, power, cooling, and operations.

You will coordinate various engineering partners through validation, review provider specifications and test results, and produce clear go/no-go recommendations on whether new capacity meets our standards. It is a high-impact, process-driven role for someone technical enough to know when a spec sheet does not add up, and additional diligence needs to be completed, and organized enough to drive many evaluations to closure in parallel. "You will deep-dive into critical hardware performance metrics, proactively identifying potential bottlenecks in cluster architecture before they impact our end customers training or inference workloads." Conduct diligence and work with engineering teams to make assessments regarding technical and operational resilience.

Responsibilities
  • Own and continuously improve the end-to-end qualification process for new compute capacity, from initial provider intake through final go/no-go recommendation.
  • Run multiple provider evaluations in parallel, setting timelines, tracking status, and keeping every stakeholder aligned on what is needed and by when.
  • Partner with infrastructure engineering, network engineering, data center engineering, and SRE teams to plan and coordinate technical validation, then translate their findings into clear decisions for leadership.
  • Review provider technical specifications and questionnaire responses for completeness and accuracy, flagging gaps, inconsistencies, and risks that warrant follow-up.
  • Conduct first-pass analysis of provider data yourself: compare specifications across suppliers , sanity-check performance claims, and surface issues before deeper engineering review.
  • Maintain the standards, templates, and documentation that define what meets spec across compute, networking, storage, power, cooling, and operational support.
  • Build a structured, auditable record of evaluation outcomes that informs sourcing decisions and scales the qualification function as the team grows.
Requirements
  • 5+ years in technical program or project management, infrastructure program management, or a comparable technical operations role, ideally involving hardware, data center, or large-scale compute environments.
  • Proven ability to run multiple complex, cross-functional workstreams to deadline, with strong organization and stakeholder management.
  • Working technical fluency across data center infrastructure: server and GPU hardware, high-performance networking (InfiniBand or Ethernet fabrics), storage, and power and cooling fundamentals; enough depth to read a detailed technical specification and know what to question.
  • Hands-on comfort with data: able to write scripts or queries (for example, Python or SQL) to compare, validate, and analyze provider specifications and test results independently.
  • Excellent written and verbal communication; able to turn dense technical detail into clear recommendations for both engineers and executives.
  • Willingness to travel to provider and data center sites as needed.
  • This is not an engineering manager role
Nice to Have
  • Experience qualifying, commissioning, or accepting GPU clusters or HPC infrastructure against defined performance and reliability standards.
  • Familiarity with AI training and inference infrastructure, including interconnect topologies, cluster bring-up, and acceptance testing.
  • Experience in AI/HPC cluster design.
  • Background working directly with hardware vendors, colocation providers, or cloud capacity providers.
About Together AI

Together AI, the AI Native Cloud, is purpose-built for AI engineers. AI application developers get high-performance inference that scales reliably, fine-tuning and reinforcement learning for creating frontier-level specialized models, and pre-training at massive scale for fully custom intelligence, all around a marketplace of leading open models that teams can run, adapt, and own. Trusted by Cursor, Decagon, ElevenLabs, Salesforce, and Zoom, Together serves 400+ trillion tokens a month.

Compensation

We offer competitive compensation, startup equity, health insurance, and other benefits, as well as flexibility in terms of remote work. The US base salary range for this full-time position is: $200-250K + equity + benefits. Our salary ranges are determined by location, level and role. Individual compensation will be determined by experience, skills, and job-related knowledge.

Equal Opportunity

Together AI is an Equal Opportunity Employer and is proud to offer equal employment opportunity to everyone regardless of race, color, ancestry, religion, sex, national origin, sexual orientation, age, citizenship, marital status, disability, gender identity, veteran status, and more. Please see our Privacy Policy at https://www.together.ai/privacy

Similar Jobs

15 Minutes Ago
In-Office or Remote
2 Locations
179K-420K Annually
Senior level
179K-420K Annually
Senior level
Artificial Intelligence • Cloud • Information Technology • Consulting
Manages HPE’s strategic relationship with ePlus, driving partner-led revenue, profitability, pipeline, forecasting, and quota attainment. Develops joint business plans, aligns HPE field sales and specialists, promotes HPE products and solutions, recruits partner resources, and builds executive-level relationships. Advises ePlus on technology trends, portfolio strategy, compliance, and customer opportunities while expanding HPE market share and partner commitment.
Top Skills: Eplus Partner EcosystemHpe Products And Services
16 Minutes Ago
Remote or Hybrid
New York, NY, USA
175K-215K Annually
Senior level
175K-215K Annually
Senior level
AdTech • Cloud • Digital Media • Information Technology • News + Entertainment • App development
Develops deep learning, computer vision, and procedural modeling algorithms for high-fidelity 2D and 3D content. Applies machine learning and computer graphics research to large media and geospatial datasets, deploys and tests systems on remote Unix machines, and manages modular code with Git. Collaborates with founders to translate product and customer needs into technical milestones.
Top Skills: Artificial Neural NetworksC++Classical Machine LearningComputer GraphicsComputer VisionDeep LearningGitGradient DescentLinear AlgebraNonlinear OptimizationNumerical MethodsProbabilityPythonStatisticsUnix Shell
17 Minutes Ago
Remote or Hybrid
New York, NY, USA
120K-165K Annually
Senior level
120K-165K Annually
Senior level
AdTech • Cloud • Digital Media • Information Technology • News + Entertainment • App development
Leads multiple cross-functional Ad Sales Technology programs from planning through launch. Owns integrated plans, dependencies, risks, budgets, milestones, governance, stakeholder communications, readiness, and change management. Partners with Product, Engineering, Data, QA, Support, and Ad Sales teams to resolve blockers and improve delivery. Develops KPI dashboards, reporting processes, retrospectives, and continuous-improvement initiatives. Requires strong program management experience in complex technology and software development environments, with familiarity with Agile delivery and tools such as Jira, Confluence, and Smartsheet.
Top Skills: AgileConfluenceDigital Ad ServingHybrid Delivery EnvironmentsJIRAExcelMS OfficeMicrosoft OutlookMicrosoft PowerpointMicrosoft WordProgrammatic AdvertisingSmartsheet

What you need to know about the NYC Tech Scene

As the undisputed financial capital of the world, New York City is an epicenter of startup funding activity. The city has a thriving fintech scene and is a major player in verticals ranging from AI to biotech, cybersecurity and digital media. It also has universities like NYU, Columbia and Cornell Tech attracting students and researchers from across the globe, providing the ecosystem with a constant influx of world-class talent. And its East Coast location and three international airports make it a perfect spot for European companies establishing a foothold in the United States.

Key Facts About NYC Tech

  • Number of Tech Workers: 549,200; 6% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Capgemini, Bloomberg, IBM, Spotify
  • Key Industries: Artificial intelligence, Fintech
  • Funding Landscape: $25.5 billion in venture capital funding in 2024 (Pitchbook)
  • Notable Investors: Greycroft, Thrive Capital, Union Square Ventures, FirstMark Capital, Tiger Global Management, Tribeca Venture Partners, Insight Partners, Two Sigma Ventures
  • Research Centers and Universities: Columbia University, New York University, Fordham University, CUNY, AI Now Institute, Flatiron Institute, C.N. Yang Institute for Theoretical Physics, NASA Space Radiation Laboratory

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account