You will act as the escalation point for our most challenging technical hurdles, ensuring that our compiler technology runs flawlessly on the world's most powerful hardware.
Think: Lead the architectural strategy for customer rollouts. You will analyze client infrastructure—evaluating power, high-speed interconnects (Infiniband/RoCE), and software environments—to plan successful cluster deployments. You will drive complex customer issues to resolution by diagnosing root causes that sit between hardware, the OS, and our application layer.
Implement:
Execute hands-on deployments of Kubernetes clusters (on-prem and cloud) tailored for GPU acceleration.
Dive deep into code and systems to detail, reproduce, and resolve issues. You will set up test environments using C#, CUDA, and ROCm to mimic customer failures.
Work directly with the latest silicon (NVIDIA H100, AMD MI300) and interconnects to ensure our software utilizes the hardware correctly.
Build:
The Knowledge Base: You will author detailed technical solutions, white papers, and "known issue" documentations. Your work will empower the rest of the team and our users to solve problems faster.
Feedback Loops: Collaborate closely with the Engineering and R&D teams. You will translate field data into clear bug reports and feature requests, helping to shape the future stability of the product.
You are a "System Doctor." You have the computer science fundamentals to understand code, but your expertise lies in making that code run reliably on physical systems.
Experience: You have a BS/MS in Computer Science, Electrical Engineering, or related field, with 8+ years of experience in system software development and hardware support. You have a proven track record in customer-facing roles.
HPC & Hardware Fluency: You have a deep understanding of GPU architectures and how they interact with the rest of the system. You are comfortable dealing with high-speed interconnects, PCIe topology, and driver stacks.
Software Ecosystem: You possess strong computer science fundamentals. You are an expert in Python and scripting for automation, but you are also comfortable navigating C#/.NET environments and the CUDA/ROCm ecosystems.
Containerization: You have practical experience deploying and debugging Kubernetes clusters in production environments.
Communication: Excellent interpersonal skills are non-negotiable. You can remain calm under pressure, communicate complex technical details to stakeholders, and manage customer expectations effectively.
Similar Jobs
What you need to know about the NYC Tech Scene
Key Facts About NYC Tech
- Number of Tech Workers: 549,200; 6% of overall workforce (2024 CompTIA survey)
- Major Tech Employers: Capgemini, Bloomberg, IBM, Spotify
- Key Industries: Artificial intelligence, Fintech
- Funding Landscape: $25.5 billion in venture capital funding in 2024 (Pitchbook)
- Notable Investors: Greycroft, Thrive Capital, Union Square Ventures, FirstMark Capital, Tiger Global Management, Tribeca Venture Partners, Insight Partners, Two Sigma Ventures
- Research Centers and Universities: Columbia University, New York University, Fordham University, CUNY, AI Now Institute, Flatiron Institute, C.N. Yang Institute for Theoretical Physics, NASA Space Radiation Laboratory



