Galaxy Logo

Galaxy

Vice President Site Reliability Engineering (Data Centers)

Posted 7 Days Ago
Remote
Hiring Remotely in USA
Senior level
Remote
Hiring Remotely in USA
Senior level
Lead an SRE automation team building and governing infrastructure-as-code, configuration management, image pipelines, and monitoring for hybrid data centers and cloud. Drive lifecycle automation, custom tooling, capacity/performance optimization, cross-team collaboration, and mentor engineers to ensure reliable, self-service infrastructure platforms.
The summary above was generated by AI

Who We Are:

Galaxy Digital Inc. (Nasdaq: GLXY) is a global leader in digital assets and data center infrastructure, growing the economy that runs on code. Galaxy delivers the onchain infrastructure that connects institutions to digital assets, including trading, advisory, asset management, staking, self-custody, and tokenization. Galaxy also develops and operates data center infrastructure to power AI and HPC workloads. Anchored by its Helios campus in Texas, Galaxy is building a multi-gigawatt pipeline of more than 5.7 GW of potential capacity, positioning it among the largest and fastest-growing data center developers in North America.

The Company is headquartered in New York City, with offices across North America, Europe, the Middle East, and Asia.

Additional information about Galaxy's businesses and products is available on www.galaxy.com.

What We Value:

We are a diverse team of free thinkers, and fast movers united to help investors and creators energize the global economy. We are looking for individuals who thrive in a culture of builders and overachievers and embrace high performance, transparent feedback, and a mission-first approach. Our culture shapes our way of working and gets us where we want to be.

  • Seek Excellence.
  • Be Selective To Be Effective.
  • Be Highly Aligned, Loosely Coupled.
  • Disagree Transparently.
  • Encourage Independent Decision-Making.
  • Build Dream Teams.

Who You Are

A collaborative and strategic leader with deep hands-on experience in Site Reliability Engineering (SRE) and infrastructure Automation. You are comfortable steering the vision for an enterprise automation roadmap while remaining technical enough to dive into the code. You treat infrastructure as a product, ensuring that your automation workflows are as reliable as the services they deploy. You have a proven track record of managing complex hybrid environments and are proactive in building self-service platforms that enhance engineering velocity and system stability.

Responsibilities

  • Automation Platform Leadership: Oversee a specialized SRE team focused on the design, deployment, and maintenance of automation toolsets as well as the systems they interact with.
  • Infrastructure as Code (IaC) Governance: Establish and enforce standards for IaC to ensure consistent, repeatable, and secure deployments across an entire infrastructure ecosystem. Strong proficiency in Terraform is required.
  • Configuration Management: Lead the strategy for automated configuration and state management, ensuring Ansible playbooks and Packer image pipelines are optimized for both Windows, Linux, and ESXi Platforms.
  • Monitoring & Observability: Manage the monitoring and health of the automation platforms themselves. Implement SLIs/SLOs to ensure the "tools that build the servers" are highly available and performant.
  • Lifecycle Management: Drive the automated lifecycle of both physical and virtual assets, from initial template creation/deployment to automated patching, scaling, and decommissioning.
  • Custom Tooling & Scripting: Lead the development of custom scripts and internal providers (Python, Go, PowerShell, Bash) to provide better insights and tooling for our systems.
  • Collaboration: Outside of the automation team you will need to be able to collaborate and foster workflows alongside the rest of the Datacenter team and be able to facilitate needs for the team as a whole.
  • Capacity & Performance: Analyze system behavior and resource utilization in virtual environments to optimize the performance of automated deployments.
  • Mentorship & Growth: Provide technical guidance and career mentorship to SREs, fostering a culture of "automate-first" and continuous improvement.

Requirements

  • 6-10 years’ experience in Infrastructure, SRE or DevOps, specifically focused on infrastructure automation at scale.
  • Deep proficiency with Terraform (providers, modules, state management) and Ansible (roles, playbooks, Tower/AWX).
  • Hands-on experience with Image Creation (i.e. Packer, Ansible, SCCM) to build standardized, hardened images for both Windows and Linux in hybrid environments.
  • Strong experience managing and automating virtual platforms such as VMware (vSphere/vCenter) as well as Cloud providers such as Azure and AWS.
  • High-level scripting skills in mediums such as Python, Go, PowerShell, and Bash.
  • Experience with observability tools (Splunk, ELK, Prometheus, or Grafana) to monitor infrastructure health and automation telemetry.
  • Good understanding of Network topology and design as well as experience with platforms such as Juniper Networks or Palo Alto.
  • Strong mastery of Git (branching strategies, PR workflows) and CI/CD platforms (Jenkins, GitLab CI, or GitHub Actions).
  • Equal comfort managing, troubleshooting, and tuning performance for both Windows Server and Linux.

Nice to Have

  • Previous work experience includes notable periods of team leadership and or management.
  • Experience with IAM platforms such as Entra ID, Active Directory, and Okta.
  • Experience with Storage solutions both block based and object based hosted either on-prem (HP Alletra, EMC, DDN) or in cloud (S3, Azure Blob).
  • Storage Backup/DR administration and management with Commvault or Veeam.

Galaxy respects diversity and seeks to provide equal employment opportunities to all employees and job applicants for employment without regard to actual or perceived age, race, color, creed, religion, sex or gender (including pregnancy, childbirth, lactation and related medical conditions), gender identity or gender expression (including transgender status), sexual orientation, marital or partnership or caregiver status, ancestry, national origin, citizenship status, disability, military or veteran status, protected medical condition as defined by applicable state or local law, genetic information or predisposing genetic characteristic, or other characteristic protected by applicable federal, state, or local laws and ordinances.

We will endeavor to make a reasonable accommodation to the known limitations of a qualified applicant with a disability unless the accommodation would impose an undue hardship on the operation of our business. If you believe you require such assistance to complete the application process or to participate in an interview, please contact [email protected]

Similar Jobs

12 Minutes Ago
Easy Apply
Remote or Hybrid
13 Locations
Easy Apply
100K-120K Annually
Senior level
100K-120K Annually
Senior level
Edtech • Kids + Family • Social Impact • Software
Lead QA for identity and account systems, drive Playwright-based test automation and API validation, shape test strategy, collaborate cross-functionally, improve reliability and release confidence, and mentor quality engineering practices across squads.
Top Skills: Api TestingClaude CodeCRMJavaScriptLmsOauthPayment ProcessorsPlaywrightSAMLSso
12 Minutes Ago
Easy Apply
Remote
United States
Easy Apply
153K-180K Annually
Senior level
153K-180K Annually
Senior level
Artificial Intelligence • Fintech • Healthtech • Software
Lead product analytics efforts by designing and evaluating experiments, building dashboards, and producing statistical insights. Partner with product, design, and engineering to measure product success, identify optimization opportunities, and embed ML-driven personalization into the product.
Top Skills: DbtHexLookerPythonSQL
14 Minutes Ago
Remote or Hybrid
Texas, USA
126K-212K Annually
Expert/Leader
126K-212K Annually
Expert/Leader
Artificial Intelligence • Cloud • Sales • Security • Software • Cybersecurity • Data Privacy
Lead a regional Healthcare sales team selling SailPoint's IGA SaaS solution to end customers and channel partners. Build pipeline and forecasting rigor, hire and develop reps, collaborate with marketing and partners, and execute 1-, 3-, and 12-month plans to drive quota attainment and long-term growth.
Top Skills: Identity SecurityIgaIga Solution SuiteSaaSSailpoint

What you need to know about the NYC Tech Scene

As the undisputed financial capital of the world, New York City is an epicenter of startup funding activity. The city has a thriving fintech scene and is a major player in verticals ranging from AI to biotech, cybersecurity and digital media. It also has universities like NYU, Columbia and Cornell Tech attracting students and researchers from across the globe, providing the ecosystem with a constant influx of world-class talent. And its East Coast location and three international airports make it a perfect spot for European companies establishing a foothold in the United States.

Key Facts About NYC Tech

  • Number of Tech Workers: 549,200; 6% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Capgemini, Bloomberg, IBM, Spotify
  • Key Industries: Artificial intelligence, Fintech
  • Funding Landscape: $25.5 billion in venture capital funding in 2024 (Pitchbook)
  • Notable Investors: Greycroft, Thrive Capital, Union Square Ventures, FirstMark Capital, Tiger Global Management, Tribeca Venture Partners, Insight Partners, Two Sigma Ventures
  • Research Centers and Universities: Columbia University, New York University, Fordham University, CUNY, AI Now Institute, Flatiron Institute, C.N. Yang Institute for Theoretical Physics, NASA Space Radiation Laboratory

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account