Andromeda (andromeda.ai) Logo

Andromeda (andromeda.ai)

Senior Site Reliability Engineer

Reposted 28 Days Ago
In-Office or Remote
Hiring Remotely in United States
Senior level
In-Office or Remote
Hiring Remotely in United States
Senior level
Design and operate large-scale GPU infrastructure for distributed AI training, ensuring reliability, performance, and efficient customer partnerships.
The summary above was generated by AI

Senior Site Reliability Engineer

Location: Global Remote / San Francisco · Full-Time

About Andromeda

Andromeda is a market and infrastructure platform to buy, sell, and operate compute.
We believe demand for compute will grow exponentially. So fast that a handful of vertically integrated providers won't be able to scale across operations, capital, supply chains, and politics to serve it. The result is a massive wave of fragmentation, with AI factories of every shape and size coming to market to fill this demand. Our job is to enable all of that fragmented compute to flow through one platform, delivering reliable capacity to model builders, research labs, and inference providers when they need it. We believe every spare electron should be made productive for AI and we're building the platform that makes that possible.


We sit at the center of three forces:

  • Companies that need reliable, high-performance compute fast

  • A fragmented global supply of GPUs across hyperscalers, neoclouds, and independent data centers

  • Capital, risk, and operational complexity that most teams are not equipped to manage

When we succeed, trillions of dollars of compute will flow through Andromeda. Builders get capacity when they need it. Providers get a reliable way to monetize, operate, and finance infrastructure at scale. Capital gets an easy way to deploy, hedge, and underwrite.
In five years, Andromeda won't just participate in the AI infrastructure market. We will shape it.

The Role

This is not a generalist SRE role.

You will design, operate, and debug large-scale GPU infrastructure used for distributed training and inference, working directly with customers pushing the limits of modern AI systems.

We’re looking for engineers who have personally run GPU clusters in production, understand the failure modes of distributed training, and can reason about performance from network fabric → kernel → framework.

What You’ll Own

  • GPU Cluster Architecture: Design and evolve multi-provider, multi-region GPU compute clusters optimized for large-scale training. Make topology-aware scheduling, networking, and storage decisions that directly impact training throughput and cost efficiency.

  • Customer Technical Partnership: Serve as the primary technical point of contact for customers running large-scale training workloads. Onboard, troubleshoot, and optimize, often in real time.

  • Reliability & Performance Engineering: Define SLOs and error budgets that account for the unique failure modes of GPU infrastructure (ECC errors, NVLink degradation, NCCL timeouts). Own capacity planning across heterogeneous GPU fleets optimized for training throughput.

  • Networking & Fabric Health: Ensure the health and performance of high-speed interconnects (InfiniBand, RoCE, NVLink) that underpin distributed training. Diagnose and resolve fabric-level issues that degrade collective operations.

  • Observability: Build deep visibility into GPU utilization, memory pressure, interconnect throughput, training job performance, and hardware health. Go well beyond standard infrastructure metrics.

  • Automation & Tooling: Build production-grade automation for cluster provisioning, GPU health checks, job scheduling, self-healing, and firmware/driver lifecycle management.

  • Incident Leadership: Lead incident response for complex, multi-layer failures spanning hardware, networking, orchestration, and ML frameworks. Drive blameless postmortems and systemic fixes.

What We’re Looking For

  • GPU Systems Expertise: Deep, hands-on experience operating large-scale GPU clusters (NVIDIA A100/H100/B200 or equivalent). You understand GPU memory hierarchies, ECC behavior, thermal throttling, and hardware failure modes from direct experience not documentation.

  • High-Performance Networking: Production experience with InfiniBand, RoCE, or NVLink fabrics in the context of distributed training. You can diagnose why an all-reduce is slow, identify a degraded link in a fat-tree topology, and reason about congestion control at scale.

  • Distributed Training & ML Frameworks: Working knowledge of how large training jobs actually run — NCCL, CUDA, PyTorch distributed, DeepSpeed, Megatron, FSDP, or similar. You don't need to write the models, but you need to understand what's happening at the systems level when a 1,000-GPU training run stalls.

  • Linux & Systems Internals: Expert-level Linux knowledge: kernel tuning, driver management (NVIDIA drivers, CUDA toolkit), cgroup/namespace internals, performance profiling at the syscall and hardware level.

  • Kubernetes & Orchestration: Strong experience running Kubernetes in production with GPU workloads, including device plugins, topology-aware scheduling, multi-cluster federation, and custom operators. Experience with Slurm or other HPC schedulers is equally valued.

  • Automation & Software Engineering: Strong engineering skills in Python, Go, or Bash. You build production-grade tools and services, not just scripts. Infrastructure-as-Code proficiency (Terraform, Helm, Ansible, or equivalent).

  • Observability & Monitoring: Hands-on experience building monitoring and alerting for GPU infrastructure, not just Prometheus/Grafana basics, but GPU-specific telemetry (DCGM, nvidia-smi, fabric manager metrics) integrated into actionable dashboards.

  • Incident Management: Proven track record leading incident response for complex distributed systems where the failure could be in hardware, firmware, networking, drivers, orchestration, or application code and you need to narrow it down fast.

Strong Candidates May Have

  • Distributed Storage: Experience with high-performance parallel file systems (VAST, Weka, Lustre, GPFS) and the checkpoint I/O and data-loading bottlenecks that come with large training runs.

  • Training Optimization: Experience profiling and optimizing distributed training performance: identifying stragglers, tuning collective communication strategies, improving MFU (Model FLOPs Utilization), and reducing idle GPU time across large runs.

  • Cluster Buildout & Hardware: Experience involved in physical cluster design - rack layout, power/cooling constraints, network topology design, and hardware validation/burn-in at scale.

  • Team Leadership: Experience leading or mentoring a team of infrastructure engineers. We're growing and need people who raise the bar for everyone around them.

Why You’ll Love It Here

This is a high-impact, senior builder’s role. You’ll have significant ownership and autonomy to shape how our systems run at a foundational level, working directly with customers and providers while architecting the infrastructure backbone for reliable, scalable AI compute. You’ll influence technical direction and help define what world-class AI infrastructure operations look like.

Andromeda Cluster is an equal opportunity employer. We celebrate diversity and are committed to creating an inclusive environment for all employees. We do not discriminate on the basis of race, religion, color, national origin, gender, sexual orientation, age, marital status, veteran status, or disability status.

Similar Jobs

Yesterday
Remote
Senior level
Senior level
Software • Automation
Owns reliability, scalability, observability, and incident response for a mission-critical SaaS platform. Responsibilities include 24x7 on-call support, root cause analysis, automation, AWS infrastructure design, EKS/Kubernetes and Docker operations, Terraform and Helm deployments, CI/CD maintenance, cloud networking, monitoring with Datadog, Grafana, and Prometheus, database operations, and migration toward Kubernetes. The role also develops self-healing systems, documentation, runbooks, and cross-functional customer-focused reliability practices.
Top Skills: AlbAmazon RdsAWSBashCi/CdDatadogDevsecopsDockerDocker SwarmEksGrafanaHelmIamInfrastructure As CodeKubernetesLinuxNlbPostgresPrometheusPythonRoute 53TerraformTerraformTransit GatewayVpcVpn
2 Days Ago
Remote
2 Locations
Senior level
Senior level
Cloud • Software
The Senior Site Reliability Engineer improves production reliability, resilience, and system availability. Responsibilities include operating Linux and Kubernetes infrastructure, automating workflows, developing tools, managing observability, responding to incidents, participating in on-call rotations, writing runbooks, and conducting blameless postmortems. The role collaborates with distributed teams and stakeholders, supports cloud infrastructure, troubleshoots live systems, and promotes DevOps and SRE best practices.
Top Skills: AWSBare-Metal InfrastructureBashCassandraCi/CdClickhouseGitGoKubernetesLinuxLogsMetricsObservabilityPostgresPythonTerraformTraces
20 Days Ago
Remote
145K-193K Annually
Senior level
145K-193K Annually
Senior level
Gaming
Own and operate infrastructure for a large-scale sports betting and media platform across cloud and production environments. Lead infrastructure migrations, build Kubernetes platform tooling and CI/CD automation, improve observability and incident response, support development teams, troubleshoot distributed systems, reduce operational toil, and mentor engineers.
Top Skills: ArgocdAWSBashCephCiliumDatadogGCPGithub ActionsGoHelmIstioKubernetesLinuxPgbouncerPostgresPythonTalos OsTerraform

What you need to know about the NYC Tech Scene

As the undisputed financial capital of the world, New York City is an epicenter of startup funding activity. The city has a thriving fintech scene and is a major player in verticals ranging from AI to biotech, cybersecurity and digital media. It also has universities like NYU, Columbia and Cornell Tech attracting students and researchers from across the globe, providing the ecosystem with a constant influx of world-class talent. And its East Coast location and three international airports make it a perfect spot for European companies establishing a foothold in the United States.

Key Facts About NYC Tech

  • Number of Tech Workers: 549,200; 6% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Capgemini, Bloomberg, IBM, Spotify
  • Key Industries: Artificial intelligence, Fintech
  • Funding Landscape: $25.5 billion in venture capital funding in 2024 (Pitchbook)
  • Notable Investors: Greycroft, Thrive Capital, Union Square Ventures, FirstMark Capital, Tiger Global Management, Tribeca Venture Partners, Insight Partners, Two Sigma Ventures
  • Research Centers and Universities: Columbia University, New York University, Fordham University, CUNY, AI Now Institute, Flatiron Institute, C.N. Yang Institute for Theoretical Physics, NASA Space Radiation Laboratory

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account