Thinking Machines Lab Logo

Thinking Machines Lab

Software Engineer, Distributed Systems

Posted Yesterday
Be an Early Applicant
In-Office
New York, NY, USA
300K-350K Annually
Senior level
In-Office
New York, NY, USA
300K-350K Annually
Senior level
Design and build fault-tolerant distributed systems for orchestration, scheduling, storage, networking, and resource allocation across large GPU and TPU clusters. Develop systems that handle hardware failures, network partitions, and large-scale workloads. Improve collective communication and performance, collaborate with research and infrastructure teams, and deliver production-quality architecture and code.
The summary above was generated by AI
About Thinking Machines

The mission of Thinking Machines is to build AI that extends human will and judgment. We are training frontier models with Inkling, developing Tinker to let people make models their own, and crafting interfaces that broaden human-AI communication. We believe the future worth building is human, and we're hiring people who want to build it.

About the Role

We're hiring a Software Engineer, Distributed Systems to design and build the core distributed systems that everything else at Thinking Machines runs on — orchestration, scheduling, storage, and networking across thousands of machines. Your work underpins both Inkling's training clusters and Tinker's serving platform, and shows up anywhere we need software to coordinate reliably at scale.

This is deep systems work. You'll be reasoning about consensus, fault tolerance, and performance under real-world failure conditions, often on problems that don't have an off-the-shelf solution.

What You'll Do
  • Design and build distributed systems for compute orchestration, scheduling, storage, and networking across large GPU and TPU clusters

  • Develop fault-tolerant systems that keep running correctly as hardware fails, networks partition, and workloads scale

  • Build the distributed storage and data orchestration layers that move and persist large volumes of training and model data

  • Improve the performance and efficiency of collective communication, scheduling, and resource allocation across thousands of machines

  • Partner with research and infrastructure teams to identify systems bottlenecks and design solutions from first principles

  • Write production-quality code and help shape the architecture of systems used company-wide

Skills & QualificationsMinimum Qualifications
  • 5+ years of experience building large-scale distributed systems

  • Proficiency in Python and Go, C++, or another systems-level language

  • Strong understanding of distributed systems fundamentals: consensus, consistency, replication, and fault tolerance

  • Experience with network programming, load balancing, or distributed storage systems

Preferred Qualifications
  • Experience building distributed compute or orchestration systems for AI or ML workloads

  • Fluent in containerization, orchestration, and distributed compute frameworks

  • Experience with specialized hardware (GPUs, TPUs) and their integration into distributed training or serving systems

  • Background at AI research labs, high-performance computing centers, or similarly demanding environments

  • Published work or open-source contributions related to distributed systems or performance engineering

  • Comfortable operating with high autonomy in a fast-changing, early-stage environment

Logistics
  • Location: This role is based in San Francisco, CA.

  • Compensation: Depending on background, skills and experience, the expected annual salary range for this position is $300,000 - $400,000 USD.

  • Visa sponsorship: We sponsor visas. While we can't guarantee success for every candidate or role, if you're the right fit, we're committed to working through the visa process together.

  • Benefits: Thinking Machines offers generous health, dental, and vision benefits, unlimited PTO, paid parental leave, and relocation support as needed.

Similar Jobs

18 Days Ago
In-Office
New York, NY, USA
320K-485K Annually
Senior level
320K-485K Annually
Senior level
Artificial Intelligence • Natural Language Processing • Generative AI
Build safety and oversight infrastructure for AI systems, including behavioral monitoring, automated enforcement, internal dashboards, multi-cloud deployments, sandboxed agent runtimes, network and data access controls, audit logging, and observability. Manage platform capacity, costs, and service-level objectives while collaborating with researchers, engineers, security, privacy, and legal teams to deploy trusted detection and mitigation systems.
Top Skills: Audit LoggingContainersDistributed SystemsInfrastructure As CodeMicrovmsMonitoring SystemsMulti-Cloud InfrastructureNetwork PolicyObservabilityPythonRustService-Level Objectives
One Month Ago
Remote or Hybrid
United States
170K-170K Annually
Senior level
170K-170K Annually
Senior level
Marketing Tech • Analytics
Design and develop low-latency, high-throughput distributed database kernels for massively parallel analytics. Responsibilities include architecture, modern C++ development, MPI/OpenMP query execution, SIMD and hardware optimization, distributed networking, database internals, in-database machine learning, vector indexing, performance tuning, code reviews, mentoring, and resolving scalability issues across large-scale clustered systems.
Top Skills: Arm NeonAvx-512B-TreesC++C++17C++20C++23CudaDpdkGcsHnswIvf-PqLsm-TreesMapreduceMpiOpenmpPurifyRdmaRoceRocmS3SimdStd::AtomicSveValgrind
One Month Ago
Remote or Hybrid
United States
250K-250K Annually
Expert/Leader
250K-250K Annually
Expert/Leader
Marketing Tech • Analytics
Design and develop massively parallel distributed database kernels for real-time analytics at terabyte and petabyte scale. Responsibilities include low-level C++ development, MPP query execution, SIMD and hardware optimization, distributed communication using MPI and RDMA technologies, database internals, vector indexing, in-database machine learning, performance tuning, architecture reviews, mentoring, and resolving scalability issues.
Top Skills: Arm NeonArm SveAvx-512C++C++17C++20C++23CudaDpdkGcsHnswIvf-PqMapreduceMpiOpenmpPurifyRdmaRoceRocmS3SimdValgrind

What you need to know about the NYC Tech Scene

As the undisputed financial capital of the world, New York City is an epicenter of startup funding activity. The city has a thriving fintech scene and is a major player in verticals ranging from AI to biotech, cybersecurity and digital media. It also has universities like NYU, Columbia and Cornell Tech attracting students and researchers from across the globe, providing the ecosystem with a constant influx of world-class talent. And its East Coast location and three international airports make it a perfect spot for European companies establishing a foothold in the United States.

Key Facts About NYC Tech

  • Number of Tech Workers: 549,200; 6% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Capgemini, Bloomberg, IBM, Spotify
  • Key Industries: Artificial intelligence, Fintech
  • Funding Landscape: $25.5 billion in venture capital funding in 2024 (Pitchbook)
  • Notable Investors: Greycroft, Thrive Capital, Union Square Ventures, FirstMark Capital, Tiger Global Management, Tribeca Venture Partners, Insight Partners, Two Sigma Ventures
  • Research Centers and Universities: Columbia University, New York University, Fordham University, CUNY, AI Now Institute, Flatiron Institute, C.N. Yang Institute for Theoretical Physics, NASA Space Radiation Laboratory

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account