Fluidstack Logo

Fluidstack

Forward Deployed Engineer, Compute Operations

Posted 2 Days Ago
Be an Early Applicant
In-Office
New York, NY, USA
224K-300K Annually
Entry level
In-Office
New York, NY, USA
224K-300K Annually
Entry level
Build fleet health, repair and RMA workflows, hardware qualification systems, facility maintenance and asset-management platforms, and structured operational procedures for large AI compute deployments. The role involves production software development, LLM and agent integrations, Kubernetes and bare-metal operations, hardware validation, incident reduction, and forward-deployed collaboration with production engineers and facility operators.
The summary above was generated by AI
About Fluidstack

We exist to make humanity more free. For most of human history, you farmed or you starved. Technology gave people more time for the things they wanted to do, instead of things they had to do. Powerful AI will be the biggest lever for human choice we've ever built - but only if models are aligned with what humanity actually wants. There are groups building AI who don't share these goals. Whoever deploys frontier compute infrastructure fastest will decide whether AI expands human freedom or shrinks it.

We're singularly focused on delivering 10 to 100s of GWs of compute faster than anyone else, rethinking every layer of the stack. We acquire power, design and build data centers, and operate them - with teams spanning hardware and software. Speed and scale are our key differentiators. Come be a part of building civilization-scale infrastructure for AI.


We hire people who care deeply about this problem space. If that is you, please apply!

How We Operate
  • Be a barrel. Full autonomy. Own things end to end, take on scope without being asked, no permission required to operate outside your core role.

  • Insane urgency. We drive everything forward as fast as possible.

  • Reason from first principles. Challenge every assumption. Zero analogy thinking, no egos, the best idea wins.

  • Love of the game. The frontier of AI is the most interesting problem of our time. We put in long hours at high intensity to push the frontier forward.

  • Build something that actually matters. If you're going to spend your time, spend it on something that matters to the world.

The Decision Team

Examples of key problems the team is working on

  • Automate the delivery of gigawatts. Every process that takes AI infrastructure from land to live compute becomes software: schedules, decisions, and todos generated from a live knowledge graph instead of chased by hand.

  • Forward-deploy beside the experts. Product teams sit with quality managers, sourcing leads, and deployment engineers on factory floors and sites, and turn their judgment into systems that reach every unit.

  • Deliver every supercomputer faster than the last. Dozens of concurrent projects feed one graph, so every lesson learned at one site becomes a preventive check at all of them.

Role Scope
  • Build the fleet health system: real-time telemetry and tiered healthchecks on every machine across Kubernetes and bare metal, rolled into one API the whole company trusts to answer "is this machine healthy," with alarms correlated into incidents that reach on-call with a drafted probable cause.

  • Turn repair and RMA into generated work: one tracked flow from failure detection through triage, parts, vendor return, and return to service, where failure thresholds route machines to repair automatically, each production engineer's shift todo list is generated for them, and time to return to service is a number the system reports.

  • Ship hardware qualification as software: burn-in, performance baselining, and new hardware validation composed into rack-level workflows, so bringing thousands of accelerators online is a repeatable run and every machine enters production with its acceptance evidence attached in the graph.

  • Run the facility on the same system as the fleet: the maintenance system for lockout tagout and work orders is live at one site and rolls out to two more, every asset register loads before the first external audit this fall, and the legacy datacenter inventory retires before the next building energizes. You own the asset model, the migration, and the day the old tools switch off.

  • Turn every runbook into a checked procedure: SOPs, training records, and technician qualifications become structured data the customer can audit, and site SLOs, deployment cycle time, and labor ramp report themselves on the dashboards a hyperscaler customer asked for. You work forward-deployed beside production engineers and facility operators, on site and on the rotation, and build what they use the next shift.

What We're Looking For

The below is a starting point. We always make space for exceptional people, so if you don't fit this role exactly, tell us where you would.

  • You've shipped production code in Go, Python, or TypeScript, and you pick up whatever language the problem demands.

  • You've built real features on LLM APIs (OpenAI, Anthropic, or open-weight models), MCP servers, and agentic frameworks.

  • You work daily with AI coding tools like Claude Code and Cursor, and you get agents doing useful work autonomously alongside you.

  • You identify problems, design the solution, and ship it without waiting for direction or approval.

  • You've moved fast under deadline while leaving foundations that other engineers extended after you moved on.

  • You've sat the on-call rotation or worked beside the people who do, and you've turned operational pain into systems that made the pager quieter.

  • Your product taste shows in what you've shipped: interfaces the engineers on the rotation call obvious, and workflows that match how the work actually happens.

  • Bonus: Production engineering or SRE on large GPU fleets. Hardware qualification or burn-in frameworks. BMC, Redfish, or IPMI tooling. CMMS, DCIM, or asset management systems. BMS/EPMS or SCADA. Prometheus and Grafana.

    Benefits:

  • Competitive total compensation package (cash + equity)

  • Health, dental, and vision insurance

  • Retirement plan

  • Generous PTO policy

We are committed to pay equity and transparency.

Fluidstack is an Equal Employment Opportunity Employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, national origin, sexual orientation, gender identity, disability and protected veterans’ status, or any other characteristic protected by law. Fluidstack will consider for employment qualified applicants with arrest and conviction records pursuant to applicable law.

You will receive a confirmation email once your application has successfully been accepted. If there is an error with your submission and you did not receive a confirmation email, please email [email protected] with your resume/CV, the role you've applied for, and the date you submitted your application-- someone from our recruiting team will be in touch.

Similar Jobs

34 Minutes Ago
Hybrid
New York, NY, USA
94K-141K Annually
Expert/Leader
94K-141K Annually
Expert/Leader
Digital Media • Information Technology • News + Entertainment
Drives revenue growth across advertiser and agency accounts in assigned industry verticals. Responsibilities include prospecting, account planning, consultative selling, solution design, proposals, negotiations, closing, forecasting, and client retention. The role develops cross-platform advertising solutions using audience insights, attribution, measurement, and data-driven strategy; presents to executives; coordinates with internal teams; and contributes to go-to-market strategies, pipeline reporting, and account performance management.
Top Skills: Addressable AdvertisingArtificial IntelligenceAttributionAudience MeasurementAutomationCross-Platform MediaProgrammatic Advertising
36 Minutes Ago
In-Office
2 Locations
153K-204K Annually
Senior level
153K-204K Annually
Senior level
Cloud • Information Technology • Machine Learning
Own the end-to-end IT SOX compliance program, including control inventories, documentation, execution, evidence collection, issue remediation, and reporting. Translate compliance requirements into scalable workflows and systems, partner with IT and finance stakeholders on SDLC and application controls, review evidence, lead root cause analyses, and report control health to leadership. The role requires deep ITGC, SOX, audit, GRC, and enterprise systems experience.
Top Skills: AuditboardCoupaIt General Controls (Itgcs)NetSuiteSalesforceSAPServicenow GrcSox Compliance WorkflowsWorkdayWorkiva
3 Hours Ago
Easy Apply
Hybrid
New York City, NY, USA
Easy Apply
0-0 Annually
Expert/Leader
0-0 Annually
Expert/Leader
Big Data • Cloud • Software • Database
Leads public-company securities compliance, SEC reporting, capital markets transactions, corporate governance, board support, M&A, strategic investments, equity programs, executive compensation, investor communications, and legal operations. Advises executives and directors, manages external counsel and budgets, and implements legal technology and AI workflows. Requires a J.D., active bar membership, and 10–12+ years of corporate legal experience spanning top-tier law firms and public technology companies.
Top Skills: Ai WorkflowsSec Reporting

What you need to know about the NYC Tech Scene

As the undisputed financial capital of the world, New York City is an epicenter of startup funding activity. The city has a thriving fintech scene and is a major player in verticals ranging from AI to biotech, cybersecurity and digital media. It also has universities like NYU, Columbia and Cornell Tech attracting students and researchers from across the globe, providing the ecosystem with a constant influx of world-class talent. And its East Coast location and three international airports make it a perfect spot for European companies establishing a foothold in the United States.

Key Facts About NYC Tech

  • Number of Tech Workers: 549,200; 6% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Capgemini, Bloomberg, IBM, Spotify
  • Key Industries: Artificial intelligence, Fintech
  • Funding Landscape: $25.5 billion in venture capital funding in 2024 (Pitchbook)
  • Notable Investors: Greycroft, Thrive Capital, Union Square Ventures, FirstMark Capital, Tiger Global Management, Tribeca Venture Partners, Insight Partners, Two Sigma Ventures
  • Research Centers and Universities: Columbia University, New York University, Fordham University, CUNY, AI Now Institute, Flatiron Institute, C.N. Yang Institute for Theoretical Physics, NASA Space Radiation Laboratory

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account