Top Tech Jobs & Startup Jobs in NYC, NY

One Month AgoSaved
Remote or Hybrid
5 Locations
272K-431K Annually
Senior level
272K-431K Annually
Senior level
Artificial Intelligence • Computer Vision • Hardware • Robotics • Metaverse
Lead NVIDIA’s RL post-training frameworks strategy and engineering ecosystem across distributed training, inference, rollout, evaluation, orchestration, and NVIDIA platforms. Build and manage globally distributed teams, prioritize upstream and internal investments, establish benchmarks and execution metrics, and drive reliable open-source integrations. Partner with research, product, hardware, CUDA, networking, and external communities to scale reinforcement learning workloads across GPUs and heterogeneous systems.
Top Skills: CudaCudnnDistributed SystemsDpoGrpoHigh-Performance ComputingKubernetesMegatron-CoreMilesMonarchNcclNemoNemo-AlignerNixlNsightNvidia GpusOpen-Source SoftwareOpenrlhfPpoRayReinforcement LearningReward ModelingRlhfSglangSkyrlSlimeSlurmTensorrt-LlmTorchtitanTransformer EngineVerl
Reposted One Month AgoSaved
In-Office
New York, NY, USA
184K-288K Annually
Senior level
184K-288K Annually
Senior level
Artificial Intelligence • Computer Vision • Hardware • Robotics • Metaverse
Develop silicon-accurate GPU kernel microbenchmarks, model-level performance analysis, and agentic kernel optimization to maximize LLM inference throughput and latency. Attribute bottlenecks across kernel, compiler, and runtime, produce optimization policies, and collaborate with compiler, hardware, kernel, and framework teams to deliver production-grade performance improvements.
Top Skills: C++CudaCuptiCutlassNcuNsysPtxPythonSassSglangTritonTrt-LlmVllm
Reposted One Month AgoSaved
In-Office or Remote
2 Locations
184K-357K Annually
Senior level
184K-357K Annually
Senior level
Artificial Intelligence • Computer Vision • Hardware • Robotics • Metaverse
Lead GPU and NVLink-based cluster design and validation for large-scale AI and HPC deployments. Advise cloud partners on architectures, perform performance modeling, debug deployment issues, support NPI rollouts, and relay field feedback to engineering.
Top Skills: Distributed TrainingHpc ClustersImexMpiNcclNmxNvidia GpusNvidia NetworkingNvlink
One Month AgoSaved
In-Office or Remote
2 Locations
184K-357K Annually
Senior level
184K-357K Annually
Senior level
Artificial Intelligence • Computer Vision • Hardware • Robotics • Metaverse
Prototype and integrate GPU-accelerated distributed data processing, compression, databases, and analytics solutions. Optimize complex data-intensive workloads across heterogeneous CPU/GPU architectures, collaborate on next-generation hardware and software designs, and work with customers and cloud service providers to deploy solutions and influence open standards. The role requires deep expertise in parallel programming, C/C++, accelerator architecture, memory systems, and compression or distributed data systems.
Top Skills: AsicAv1CC++CpuCudaGpuH.264H.265MetalMpiNicsNpuOpenaccOpenmpProresPthreadsRocmStorage I/OTbb
Reposted One Month AgoSaved
In-Office
New York, NY, USA
152K-288K Annually
Senior level
152K-288K Annually
Senior level
Artificial Intelligence • Computer Vision • Hardware • Robotics • Metaverse
As a Senior AI Developer Technology Engineer at NVIDIA, you'll design and optimize AI and HPC workloads, solve performance issues, and influence future hardware and software design.
Top Skills: C/C++CudaCutileMpiOpenaccOpenmpPthreadsTbbTensorrtTensorrt-Llm
New

Track Smarter, Apply Better.

Ditch the spreadsheets. Organize your job search with our freeApplication Tracker.

Use For Free
Application Tracker Preview
One Month AgoSaved
In-Office or Remote
5 Locations
184K-357K Annually
Senior level
184K-357K Annually
Senior level
Artificial Intelligence • Computer Vision • Hardware • Robotics • Metaverse
Design and optimize high-performance C++/CUDA libraries for GPU-based data processing. Improve multi-node GPU communication (UcxExchange), optimize Presto GPU query performance, and lead enhancements across Presto, Velox, and cuDF projects for efficient DataFrame and database acceleration.
Top Skills: C++CudaCudfGpuLibcudfMulti-Node GpuNcclPrestoRapidsUcxVelox
Reposted One Month AgoSaved
In-Office or Remote
5 Locations
184K-288K Annually
Senior level
184K-288K Annually
Senior level
Artificial Intelligence • Computer Vision • Hardware • Robotics • Metaverse
As a Senior Solutions Architect, you will assist customers in building solutions with NVIDIA's AI technology, focusing on Generative AI and Large Language Models while collaborating across teams for performance analysis.
Top Skills: AIDeep LearningDockerDynamoGpuKubernetesLlmNvidia NimPythonPyTorchTensorFlowTensorrt
Reposted One Month AgoSaved
Remote or Hybrid
6 Locations
152K-288K Annually
Senior level
152K-288K Annually
Senior level
Artificial Intelligence • Computer Vision • Hardware • Robotics • Metaverse
Design, build, and optimize GPU‑accelerated inference software and open‑source frameworks (vLLM, SGLang, FlashInfer). Improve performance and scale for LLMs and generative models across NVIDIA GPUs and edge SoCs using CUDA, CUTLASS, Triton, NCCL and profiling tools. Collaborate across teams to enable efficient model serving and deployment.
Top Skills: CC++CudaCuda KernelsCutlassFlashinferNcclNvshmemOai TritonPythonPyTorchSglangTritonVllm
Reposted One Month AgoSaved
In-Office or Remote
5 Locations
148K-259K Annually
Senior level
148K-259K Annually
Senior level
Artificial Intelligence • Computer Vision • Hardware • Robotics • Metaverse
Create and deliver developer-focused technical content and workshops that teach CUDA and NVIDIA GPU technologies. Evaluate new CUDA features, produce tutorials and blog posts, capture developer feedback to inform product decisions, and collaborate across product, engineering, and marketing teams to drive developer adoption.
Top Skills: C++CudaNvidia GpuPython
Reposted One Month AgoSaved
In-Office or Remote
4 Locations
184K-288K Annually
Senior level
184K-288K Annually
Senior level
Artificial Intelligence • Computer Vision • Hardware • Robotics • Metaverse
Advise ISVs on designing, deploying, and optimizing large-scale accelerated AI infrastructure. Lead architecture reviews, POCs, benchmarks, and production deployment guidance across compute, networking, storage, containers, orchestration, observability, and CI/CD. Create reference architectures, sizing guidance, technical playbooks, demos, and whitepapers. Support cluster monitoring, reliability, and performance improvements. Up to 20% travel for customer engagements.
Top Skills: Ci/CdContainersDistributed Training FrameworksEthernetGpuHplInfinibandMlperfNcclObservabilityOpenmpiOrchestrationRdmaSchedulingTelemetry
All Filters
JobType
New Jobs
Job Category
Experience
Industry
Company Name
Company Size

Sign up now Access later

Create Free Account