Andromeda (andromeda.ai) Logo

Andromeda (andromeda.ai)

HPC Architect

Posted Yesterday
Be an Early Applicant
In-Office or Remote
Hiring Remotely in USA
Senior level
In-Office or Remote
Hiring Remotely in USA
Senior level
Own technical qualification of compute providers: define acceptance standards and benchmarks, validate clusters (fabric, GPUs, storage, orchestration), guide provider remediation, and maintain technical relationships to ensure reliable production performance.
The summary above was generated by AI
HPC ArchitectLocation: North America Remote/SF-Hybrid · Full-Time

About Andromeda

Andromeda Cluster was founded by Nat Friedman and Daniel Gross to give early-stage startups access to the kind of scaled AI infrastructure once reserved only for hyperscalers.

We began with a single managed cluster — but it filled almost instantly. Since then, we’ve been quietly building the systems, network, and orchestration layer that makes the world’s AI infrastructure more accessible.

Today, Andromeda works with leading AI labs, data centers, and cloud providers to deliver compute when and where it’s needed most. Our platform routes training and inference jobs across global supply, unlocking flexibility and efficiency in one of the fastest-growing markets on earth.

Our long-term vision is to build the liquidity layer for global AI compute. We are expanding to new frontiers to find the brightest that work in AI infrastructure, research and engineering.

The Role

This role owns our provider relationships on the technical side. You're the person who decides whether a provider's cluster is good enough to join the Andromeda network, and the person who helps them get there when it isn't. That means vetting prospective providers against our quality bar and then working alongside their engineers to bring their clusters onto the network cleanly.

You'll work closely with our compute procurement team to identify and qualify new providers, and you'll be the standing technical relationship with the providers we already have. When procurement finds a promising provider, you're the one who validates the claims. When a provider says their fabric is non-blocking and their nodes are burn-in tested, you're the one who verifies it by reviewing their evidence, and sometimes by running the tests yourself on their hardware.

The qualification bar you'll hold providers to mostly doesn't exist yet in written form. Some of it lives informally in how our SREs evaluate clusters today; much of it hasn't been defined at all. You'd be the first person in this role, so a large part of the job is building that bar: the acceptance test suite, the quality thresholds, and the technical standards a provider must meet. Then making them rigorous enough that we can trust them and clear enough that providers can build to them.

This role is the technical counterpart to our Provider Technical Program Manager, who owns delivery timelines, escalations, and incident command. You own the technical judgment: is this cluster ready, what's wrong with it, and what will it take to fix. During provider incidents, command and coordination sit with the program manager and technical response sits with our SREs. Your involvement is upstream, making sure clusters that would have caused those incidents never make it onto the network.

What You’ll Do

  • Vet prospective compute providers: assess cluster architecture, GPU hardware, network fabric, storage, and orchestration against Andromeda's quality metrics, and make the call on whether a cluster qualifies.

  • Define the qualification bar itself. Build the acceptance test suite, benchmark methodology, and quality thresholds from scratch, formalizing what currently exists only as SRE tribal knowledge. Own and evolve these standards as the fleet and the market change.

  • Run validation hands-on where it matters: burn-in testing, fabric validation (InfiniBand/RoCE), NCCL and application-level benchmarks, storage performance testing. For routine or repeat validation, define the methodology and review results rather than executing everything yourself.

  • Guide providers through technical onboarding: work directly with their data-center and platform engineers to remediate gaps, tune configurations, and bring clusters up to the bar on a predictable path.

  • Partner with compute procurement to identify and qualify new providers. Perform technical due diligence during sourcing, and a clear read on how much remediation a candidate cluster needs before procurement commits.

  • Build and maintain technical relationships with existing providers: you're the engineer their engineers call, and the one who spots architectural or quality drift before it becomes a customer problem.

  • Feed what you learn back into provider-facing standards and internal documentation, so each qualification is faster and more consistent than the last.

What We’re Looking For

  • Deep HPC experience: you've designed, built, or operated GPU clusters at meaningful scale and understand what makes them fast, stable, and debuggable.

  • Strong fabric knowledge, ideally both InfiniBand and RoCE. You can evaluate a topology, interpret fabric diagnostics, and identify why a fabric underperforms, not just that it does.

  • Experience with distributed orchestration and the HPC software stack: Slurm, Kubernetes, OpenMPI or equivalent, and the Linux systems engineering underneath all of it.

  • Data-center literacy: power, cooling, cabling, and physical-layer realities. You can walk a provider's facility and know what questions to ask.

  • Benchmarking judgment. You know which numbers matter for large-scale training workloads, how providers game them, and how to design tests that can't be gamed.

  • The ability to write standards others can build against: precise, testable, and usable by a provider's engineering team without you in the room.

  • Credibility in the room with external engineering teams, including the ability to deliver a failing grade to a provider who wants your business, and keep the relationship intact.

  • Comfort with ambiguity. The qualification framework you'll apply mostly doesn't exist yet; you'll write it.

Strong Candidates May Have

  • Experience inside a neocloud, hyperscaler, colocation, or data-center provider.

  • NVIDIA data-center GPU depth: DGX/HGX platforms, NVLink/NVSwitch, GPU health and RAS behavior at fleet scale.

  • Experience supporting AI research labs or other large-scale training customers, and familiarity with what their workloads punish: stragglers, fabric jitter, storage stalls.

  • Background in cluster acceptance testing or site bring-up. Ideally you've taken a cluster from delivery to production before.

What Success Looks Like

Within your first year, you'll have delivered:

  • A qualification bar that's written down and trusted. Andromeda's acceptance criteria, benchmark suite, and quality thresholds exist as documents and tooling — not tribal knowledge. SREs, procurement, and providers all work from the same bar, and a passing grade from you means the cluster performs in production.

  • Providers who build to our standards before we ever test them. Prospective providers get our technical requirements upfront and arrive closer to qualified, because the standards are clear enough to engineer against. Time-from-sourcing-to-qualified drops with each provider.

  • Onboarding that's technically boring. Clusters that join the network have been validated the same way every time, and the surprises that used to appear in a customer's training run get caught in acceptance instead.

  • A technical relationship providers value. Provider engineering teams treat you as the authority on what Andromeda needs and the first call when they're planning a build-out so that we hear about architecture decisions early enough to influence them.

  • A procurement partnership that changes what we buy. Procurement's provider pipeline reflects your technical due diligence, and deals are shaped by a realistic view of remediation cost.

Why You’ll Love It Here

  • High-growth environment: Get in early at a company at the center of the AI infrastructure boom

  • Ownership: First HPC Architect for the solutions engineering team, you’ll get to build this function from the ground up

  • Competitive compensation: + meaningful equity

  • Comprehensive benefits: for you and your dependents, including healthcare, dental, and vision coverage, 401(k), and unlimited PTO

Andromeda Cluster is an equal opportunity employer. We celebrate diversity and are committed to creating an inclusive environment for all employees. We do not discriminate on the basis of race, religion, color, national origin, gender, sexual orientation, age, marital status, veteran status, or disability status.

Similar Jobs

23 Days Ago
In-Office or Remote
161K-390K Annually
Expert/Leader
161K-390K Annually
Expert/Leader
Artificial Intelligence • Cloud • Information Technology • Consulting
This role involves developing strategies for electrical hardware design, consulting on new product development, mentoring junior staff, and leading innovative projects in HPC and AI systems.
Top Skills: AIAmdCray SlingshotCudaEthernetHpcInfinibandNvidiaRocm
14 Hours Ago
Remote or Hybrid
Senior level
Senior level
AdTech • Cloud • Digital Media • Information Technology • News + Entertainment • App development
Lead design and implement major gameplay systems from concept through prototype to polish using Unreal Engine 5 and Blueprint. Prototype, optimize, and scale systemic/open-world mechanics, collaborate with engineering/art/UX, drive iteration with playtests and telemetry, mentor designers, document technical specifications, and mitigate production/technical risks.
Top Skills: Ai BehaviorsBehavior TreesBlueprintC#Editor Tools/Editor UtilitiesLuaMultiplayer/Network-Aware SystemsNavigationPerformance ProfilingPythonState TreesTelemetryUnreal Engine 5
16 Hours Ago
Remote or Hybrid
Senior level
Senior level
Cloud • Information Technology • Security • Software • Cybersecurity
As a Senior Solutions Engineer, you'll engage with enterprise customers, addressing technical sales, creating presentations, conducting demos, and leading proof of concepts for Cloudflare solutions.
Top Skills: AWSAzureCdnCloudflareGCPNetworkingSaaSSecuritySIEMSoftware Distribution

What you need to know about the NYC Tech Scene

As the undisputed financial capital of the world, New York City is an epicenter of startup funding activity. The city has a thriving fintech scene and is a major player in verticals ranging from AI to biotech, cybersecurity and digital media. It also has universities like NYU, Columbia and Cornell Tech attracting students and researchers from across the globe, providing the ecosystem with a constant influx of world-class talent. And its East Coast location and three international airports make it a perfect spot for European companies establishing a foothold in the United States.

Key Facts About NYC Tech

  • Number of Tech Workers: 549,200; 6% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Capgemini, Bloomberg, IBM, Spotify
  • Key Industries: Artificial intelligence, Fintech
  • Funding Landscape: $25.5 billion in venture capital funding in 2024 (Pitchbook)
  • Notable Investors: Greycroft, Thrive Capital, Union Square Ventures, FirstMark Capital, Tiger Global Management, Tribeca Venture Partners, Insight Partners, Two Sigma Ventures
  • Research Centers and Universities: Columbia University, New York University, Fordham University, CUNY, AI Now Institute, Flatiron Institute, C.N. Yang Institute for Theoretical Physics, NASA Space Radiation Laboratory

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account