NVIDIA Logo

NVIDIA

Senior Platform Engineer, Network Infrastructure - DGX Cloud

Reposted One Month Ago
In-Office or Remote
4 Locations
168K-334K Annually
Senior level
In-Office or Remote
4 Locations
168K-334K Annually
Senior level
Own design, automation, and lifecycle of Kubernetes platform for global network infrastructure. Build production-quality tooling for cluster provisioning, upgrades, GitOps delivery, observability, and recovery. Provide production support and incident response for network services on the platform, drive root-cause analysis, and establish platform standards and runbooks.
The summary above was generated by AI

Cloud Foundations Reliability (CFR) is part of NVIDIA’s Global Network Infrastructure (GNI) organization. We deploy, integrate, and operate the Kubernetes-based platform and shared services used to provision, monitor, and operate NVIDIA’s global network across data centers, colocation facilities, and cloud environments. The team owns the architecture and lifecycle of this platform, including cluster provisioning and upgrades, GitOps delivery, observability, capacity, and service enablement. We build software and automation to standardize how network platforms and services are deployed, scaled, and managed across environments.

We are looking for a hands-on senior engineer to own the lifecycle and automation of the Kubernetes platform supporting GNI network systems. You will also provide production support for network services running on the platform, partnering with their engineering owners when issues or changes cross the platform boundary. You will take complex problems from design through production and remain accountable for the outcome. You will bring deep Kubernetes expertise and help establish consistent engineering practices across the US and Bangalore teams. This is a senior individual contributor role with end-to-end ownership and production responsibility.

What You’ll Be Doing:

  • Design, build, and operate the Kubernetes platform that powers GNI network automation, telemetry, and operations across data center, colocation, and cloud environments.

  • Own the lifecycle management for GNI Kubernetes environments, including cluster onboarding, upgrades, capacity, availability, and recovery.

  • Develop production-quality software and automation for cluster provisioning, validation, upgrades, remediation, and safe multi-cluster delivery through GitOps.

  • Provide production support for network services hosted on the platform, working with Network Automation and service teams that retain ownership of application architecture, code, and features.

  • Diagnose complex Kubernetes platform and hosted-service failures involving control-plane health, cluster networking, storage, scheduling, workload placement, and multi-cluster dependencies. Drive issues from initial signal through verified resolution.

  • Define production-readiness and observability standards for the platform and hosted network services, including health signals, capacity, alerts, runbooks, and recovery.

  • Participate in CFR’s production on-call rotation, including scheduled after-hours and weekend coverage. Lead incident response and recovery, then drive corrective actions to completion.

What We Need to See:

  • Bachelor’s degree in Computer Science, Engineering, or a related field, or equivalent experience.

  • 8+ years of experience building or operating production Kubernetes platforms, network infrastructure, or distributed systems.

  • Deep experience with Kubernetes at scale, including cluster lifecycle, upgrades, networking, storage, and recovery.

  • Proficiency in at least one general-purpose programming language, such as Go or Python.

  • Experience with GitOps, infrastructure as code, CI/CD, and automated production delivery.

  • Experience deploying and supporting network automation or telemetry services on Kubernetes.

  • Experience with production on-call, incident response, root-cause analysis, and driving corrective actions to completion.

Ways to Stand Out From the Crowd:

  • Strong knowledge of IP routing, data center fabrics, and cloud networking is a great plus.

  • Experience designing and operating large, multi-region Kubernetes fleets, including fleet-wide upgrades and recovery.

  • Hands-on experience with Cluster API (CAPI) and Metal3 for bare-metal provisioning, cluster lifecycle, machine remediation, and upgrades.

  • Experience building Kubernetes controllers or operators in Go using custom resources and reconciliation patterns. Experience designing or operating network automation and telemetry services on Kubernetes at global scale.

  • Contributions to Cluster API, Metal3, or other open-source Kubernetes infrastructure projects.

NVIDIA’s deep learning platforms have made major impact to various fields is broadly used across leading academic institutions, start-ups, and industry, including the world’s largest Internet companies. We need passionate, hard-working and creative people to help us take on more of these unique opportunities in deep learning cloud solutions. NVIDIA is widely considered to be one of the technology world’s most desirable employers. We have some of the most forward-thinking and hard-working people in the world working for us. Are you creative and autonomous? Do you love a challenge? If so, we want to hear from you.

Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 176,000 USD - 276,000 USD for Level 4, and 208,000 USD - 333,500 USD for Level 5.

You will also be eligible for equity and benefits.

Applications for this job will be accepted at least until September 14, 2026.

This posting is for an existing vacancy. 

NVIDIA uses AI tools in its recruiting processes.

NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.

Similar Jobs

50 Minutes Ago
Remote or Hybrid
US
124K-170K Annually
Mid level
124K-170K Annually
Mid level
Information Technology
Owns the strategy, roadmap, backlog, and delivery of professional services products. Partners with technical and business stakeholders to define product vision, prioritize features, manage Agile release planning, track KPIs, conduct research, and communicate product direction to senior leaders. The role requires strong analytical, presentation, negotiation, and cross-functional collaboration skills, plus experience with professional or managed services and Agile methodologies.
Top Skills: AgileCertinia PsaKanbanLucidchartMS OfficePowerPointSafeSalesforce CpqSalesforce CRMScrum
50 Minutes Ago
Remote or Hybrid
US
132K-193K Annually
Expert/Leader
132K-193K Annually
Expert/Leader
Information Technology
Leads strategic technology consulting engagements and advises clients on transformation, risk, and business objectives. Defines and governs enterprise integration architectures across ServiceNow and business systems, using APIs, workflows, data architecture, and security controls. Develops recommendations, roadmaps, executive presentations, and architectural standards; manages stakeholders, project governance, business development, and client relationships. Mentors consultants and designs secure, scalable, maintainable integration solutions across ServiceNow modules and enterprise technologies.
Top Skills: AIAnsibleAPIsArtifactoryCommvaultCyberarkData ArchitectureDigicertGithub Ci/CdHashicorp VaultKubernetesPanoramaSecurity ControlsServicenowWorkflow OrchestrationXsoar
5 Hours Ago
Remote or Hybrid
90K-138K Annually
Senior level
90K-138K Annually
Senior level
Automotive • Professional Services • Software • Consulting • Energy • Chemical • Renewable Energy
Leads and develops a team of Sales Executives to achieve revenue, quota, pipeline, forecasting, and growth objectives. Responsibilities include coaching, performance management, territory and account strategy, prospecting, cross-selling, CRM data integrity, customer retention, recruiting, onboarding, and collaboration with internal teams. The role also manages escalations, ensures sales-process compliance, and drives regional and business-unit objectives.
Top Skills: Salesforce CRM

What you need to know about the NYC Tech Scene

As the undisputed financial capital of the world, New York City is an epicenter of startup funding activity. The city has a thriving fintech scene and is a major player in verticals ranging from AI to biotech, cybersecurity and digital media. It also has universities like NYU, Columbia and Cornell Tech attracting students and researchers from across the globe, providing the ecosystem with a constant influx of world-class talent. And its East Coast location and three international airports make it a perfect spot for European companies establishing a foothold in the United States.

Key Facts About NYC Tech

  • Number of Tech Workers: 549,200; 6% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Capgemini, Bloomberg, IBM, Spotify
  • Key Industries: Artificial intelligence, Fintech
  • Funding Landscape: $25.5 billion in venture capital funding in 2024 (Pitchbook)
  • Notable Investors: Greycroft, Thrive Capital, Union Square Ventures, FirstMark Capital, Tiger Global Management, Tribeca Venture Partners, Insight Partners, Two Sigma Ventures
  • Research Centers and Universities: Columbia University, New York University, Fordham University, CUNY, AI Now Institute, Flatiron Institute, C.N. Yang Institute for Theoretical Physics, NASA Space Radiation Laboratory

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account