Thinking Machines Lab Logo

Thinking Machines Lab

Site Reliability Engineer, Post Training

Posted 5 Days Ago
In-Office
New York, NY, USA
350K-475K Annually
Mid level
In-Office
New York, NY, USA
350K-475K Annually
Mid level
Own the reliability, performance, and uptime of large-scale post-training and reinforcement learning systems. Debug distributed failures across accelerators, networking, storage, schedulers, and training frameworks; build monitoring, alerting, recovery, checkpointing, and scheduling tools; improve cluster utilization and fault tolerance; support production model runs through on-call rotations and postmortems; and partner closely with research teams during active training runs.
The summary above was generated by AI
About Thinking Machines

The mission of Thinking Machines is to build AI that extends human will and judgment. We are training frontier models with Inkling, developing Tinker to let people make models their own, and crafting interfaces that broaden human-AI communication. We believe the future worth building is human, and we're hiring people who want to build it.

About the Role

We're hiring a Site Reliability Engineer (SRE) to keep our post-training and reinforcement learning (RL) systems fast, reliable, and easy for researchers to iterate on. Think of this as a production engineering or site reliability role built around model training: you'll own the health of the training runs, clusters, and pipelines that power post-training and RL at Thinking Machines.

You'll work side by side with research teams during active model runs — debugging failures in real time, hardening infrastructure against the next class of problem, and building the tooling and automation that let researchers spend their time on the science instead of babysitting jobs. This role has real ownership: you'll be the person a research team calls when a run stalls at 2am, and the person who makes sure it doesn't happen again.

What You’ll Do
  • Own the reliability, performance, and uptime of large-scale post-training and RL training jobs, from launch through completion

  • Partner directly with research teams during active model runs, embedding with them to unblock training and speed up iteration

  • Debug failures across the full stack — accelerators, networking, storage, schedulers, and training frameworks — and drive issues to root cause

  • Build monitoring, alerting, and automated recovery so runs self-heal or fail fast instead of silently stalling

  • Improve checkpointing, fault tolerance, and job scheduling so hardware failures cost minutes, not days of compute

  • Build internal tools that reduce toil and improve cluster utilization across post-training and RL workloads

  • Participate in an on-call rotation supporting production model runs

  • Write postmortems and turn recurring failure patterns into permanent infrastructure fixes

Skills & QualificationsMinimum Qualifications
  • 4+ years of experience as a production engineer, site reliability engineer, or infrastructure engineer operating large-scale distributed systems in production

  • Track record debugging complex failures across distributed systems — networking, hardware, kernel, or scheduler issues

  • Strong software engineering skills in Python and/or Go/C++, with the judgment to know when to script a fix versus build a system

  • Solid grounding in Linux systems internals and networking fundamentals

  • Comfortable owning production systems, including participating in on-call rotations

Preferred Qualifications
  • Experience operating GPU or TPU training clusters at scale

  • Familiarity with post-training and RL techniques (e.g., RLHF, PPO, DPO) and the infrastructure challenges specific to them, such as reward model serving, rollout generation, and mixed training/inference workloads

  • Experience with distributed training frameworks (e.g., PyTorch, Ray) and job schedulers (e.g., Slurm, Kubernetes)

  • Experience with high-performance networking (e.g., InfiniBand, RDMA, NCCL) and its role in distributed training performance

  • Experience building observability tooling purpose-built for ML training, not just general infrastructure

  • A track record of thriving in fast-changing, research-driven environments where priorities shift with the science

Logistics
  • Location: This role is based in San Francisco, CA.

  • Compensation: Depending on background, skills and experience, the expected annual salary range for this position is $350,000 - $475,000 USD.

  • Visa sponsorship: We sponsor visas. While we can't guarantee success for every candidate or role, if you're the right fit, we're committed to working through the visa process together.

  • Benefits: Thinking Machines offers generous health, dental, and vision benefits, unlimited PTO, paid parental leave, and relocation support as needed.

Similar Jobs

An Hour Ago
Hybrid
New York, NY, USA
Senior level
Senior level
Financial Services
Develop and scale customer journeys for offers and shopping, translating customer insights and product signals into audience strategies, messaging, campaign concepts, creative briefs, experimentation roadmaps, and optimization recommendations. Partner with analytics, product, creative, marketing technology, and campaign teams to launch initiatives, measure performance, improve activation and redemption, and manage partner campaign workflows and governance.
An Hour Ago
Hybrid
204K-325K Annually
Senior level
204K-325K Annually
Senior level
Blockchain • Fintech • Payments • Consulting • Cryptocurrency • Cybersecurity • Quantum Computing
Lead global regulatory strategy and execution for technology infrastructure, operational resilience, and enterprise resilience. Interpret regulations, translate them into technology control frameworks and operational guidance, advise on third-party technology contracts, coordinate cross-functional implementation, and engage regulators and legal stakeholders across jurisdictions. The role requires a law degree and significant regulatory experience in payments, banking, or financial services.
An Hour Ago
Hybrid
New York, NY, USA
245K-391K Annually
Expert/Leader
245K-391K Annually
Expert/Leader
Blockchain • Fintech • Payments • Consulting • Cryptocurrency • Cybersecurity • Quantum Computing
Leads Mastercard’s token authentication solutions domain, owning strategy, roadmaps, product development, deployment, ecosystem partnerships, and scalable authentication services for digital payments. The role oversees payment passkeys, enhanced data sharing, and Digital Payment Credentials across tokenized transactions. Responsibilities include executive alignment, business cases, engineering collaboration, market adoption, standards engagement, and building a global high-performing product organization.
Top Skills: APIsDigital Payment CredentialsFidoPayment PasskeysTokenizationW3C

What you need to know about the NYC Tech Scene

As the undisputed financial capital of the world, New York City is an epicenter of startup funding activity. The city has a thriving fintech scene and is a major player in verticals ranging from AI to biotech, cybersecurity and digital media. It also has universities like NYU, Columbia and Cornell Tech attracting students and researchers from across the globe, providing the ecosystem with a constant influx of world-class talent. And its East Coast location and three international airports make it a perfect spot for European companies establishing a foothold in the United States.

Key Facts About NYC Tech

  • Number of Tech Workers: 549,200; 6% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Capgemini, Bloomberg, IBM, Spotify
  • Key Industries: Artificial intelligence, Fintech
  • Funding Landscape: $25.5 billion in venture capital funding in 2024 (Pitchbook)
  • Notable Investors: Greycroft, Thrive Capital, Union Square Ventures, FirstMark Capital, Tiger Global Management, Tribeca Venture Partners, Insight Partners, Two Sigma Ventures
  • Research Centers and Universities: Columbia University, New York University, Fordham University, CUNY, AI Now Institute, Flatiron Institute, C.N. Yang Institute for Theoretical Physics, NASA Space Radiation Laboratory

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account