NVIDIA Logo

NVIDIA

Senior Data and Platform Engineer

Posted An Hour Ago
In-Office or Remote
4 Locations
140K-270K Annually
Senior level
In-Office or Remote
4 Locations
140K-270K Annually
Senior level
Build and operate NVIDIA DGX Cloud’s data platform, including batch and streaming pipelines, distributed data workloads, data products, quality monitoring, security, observability, and self-service consumption. Own systems from architecture through deployment and incident response. Improve platform reliability, scalability, and engineering practices while mentoring teammates. Work primarily with Python and SQL across cloud infrastructure, databases, Spark, telemetry, and operational data.
The summary above was generated by AI

NVIDIA’s DGX Cloud organization is seeking a Senior Data Engineer to become part of its data team! We develop the reliable data foundation that supports fleet health, capacity, utilization, cost, reliability, and operational decision-making throughout DGX Cloud. Our platform supports engineering, operations, finance, and product teams managing and expanding large GPU fleets across cloud service providers and NVIDIA Cloud Partners. We are looking for a practical engineer and technical lead to take charge of a key part of the Navigator data platform. We develop the systems that transform distributed infrastructure telemetry and operational data into dependable, managed data products that support fleet health, capacity, utilization, cost, and operational decisions.

We are seeking a hands-on, platform-minded engineer to build and evolve the systems that turn distributed infrastructure telemetry and operational data into reliable, governed data products. You will work across ingestion, transformation, data quality, platform architecture, security, observability, and self-service consumption to help make Navigator and the DGXC data platform a dependable source of truth. We do expect strong engineering fundamentals, experience operating production systems, and the ability to learn new platforms and domains quickly.

What you'll be doing:

  • Own systems end to end. For example, work from ambiguous customer and operational needs through architecture, implementation, deployment, observability, incident response, and ongoing support.

  • Construct data pipelines and products. Such as designing and maintain batch and streaming ingestion, transformation, reconciliation, and serving paths for fleet, capacity, utilization, cost, scheduling, and operational telemetry.

  • Build shared libraries, workflow and DAG or equivalent experience abstractions to evolve the data platform. Develop deployment tooling, data contracts, and paved-road patterns that improve team speed and safety.

  • Engineer reliable distributed workloads. As well as diagnose correctness and performance issues across applications, SQL engines, Spark jobs, storage systems, networks, and cloud services. Build for retries, idempotency, backfills, schema evolution, and partial failure.

  • Treat security as part of the build. For example, applying least privilege, service identities, secrets management, access controls, environment isolation, auditability, and safe operational practices throughout the system lifecycle.

  • ​Improve quality and operations: Establish automated tests, data-quality checks, lineage, freshness and completeness monitoring, actionable alerting, SLOs, and clear ownership.

  • Deliver consumption experiences. Such as making trusted data usable through well-modeled tables, APIs, automation, dashboards, and focused internal applications—not only through one-off queries.

  • Raise the engineering bar. Lead build reviews, communicate tradeoffs, mentor other engineers, and improve the team's architecture, testing, debugging, and operational practices.

What we need to see:

  • BS or MS in Computer Science, Engineering, or a related field, or equivalent experience.

  • 5+ years of experience building and operating production software, data platforms, backend infrastructure, databases, or distributed systems.

  • Strong software-engineering fundamentals and production proficiency in Python or another backend or systems language, with the ability and willingness to work primarily in Python and SQL.

  • Deep hands-on experience in at least one of the following areas: Distributed data processing using Spark or a comparable compute framework, Relational, distributed, or analytical database architecture and operation at scale, Production ETL, change-data-capture, streaming, or event-processing systems, Backend or cloud-platform systems that process, transform, or serve substantial data volumes, Strong SQL and data-modeling skills, including a practical understanding of query performance, schema evolution, incremental processing, consistency, and analytical consumption patterns.

  • Demonstrated ability to debug unfamiliar systems across multiple layers using logs, metrics, traces, query plans, profiles, and controlled experiments to find root causes.

  • Experience operating services or pipelines in a cloud or similarly complex production environment, including testing, CI/CD, monitoring, alerting, rollback, and incident response.

  • Working knowledge of secure platform development, including identity and access management, least privilege, secret handling, trust boundaries, and safe multi-environment deployments.

  • Ability to make sound architectural tradeoffs, own work through ambiguity, and communicate effectively with users, partner teams, and engineers from different fields.

  • A track record of learning unfamiliar technologies and domains and turning that learning into maintainable systems and reusable team practices.

  • Experience with AI agents and LLM-supported workflow automation, particularly as applied to engineering and operational activities.

Ways to stand out from the crowd:

  • Experience with Databricks, Apache Spark, PySpark, Spark SQL, Delta Lake, Unity Catalog, or another modern lakehouse or distributed-compute platform.

  • Experience with Kafka or another streaming platform, change-data capture, event development, partitioning, consumer groups, offset management, or other high-volume event systems.

  • Experience with scaling, migrating, or performance-tuning relational, distributed, time-series, object-storage, or search-focused data systems, including Elasticsearch or OpenSearch.

  • Background working with AWS, Azure, GCP, Kubernetes, Slurm, compute clusters, GPU-accelerated infrastructure, or fleet-scale telemetry.

  • Experience developing agentic systems, LLM-enabled workflow automation, harness engineering, or dependable evaluation and operational tooling for AI agents.

Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 140,000 USD - 224,250 USD for Level 3, and 168,000 USD - 270,250 USD for Level 4.

You will also be eligible for equity and benefits.

Applications for this job will be accepted at least until September 25, 2026.

This posting is for an existing vacancy. 

NVIDIA uses AI tools in its recruiting processes.

NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.

Similar Jobs

25 Days Ago
Easy Apply
Remote
United States
Easy Apply
130K-160K Annually
Senior level
130K-160K Annually
Senior level
Healthtech • Software
Designs and delivers large-scale, governed data pipelines and infrastructure across the full data lifecycle. Responsibilities include coding in Python and SQL, optimizing Airflow, dbt, AWS, Kafka, and Iceberg/Parquet workflows, implementing data quality and observability, contributing to architecture and tool evaluations, mentoring engineers, collaborating with stakeholders, and participating in on-call support.
Top Skills: AirflowApache IcebergAPIsAthenaAWSCi/CdDatabricksDbtEmrHipaaKafkaParquetPythonS3SnowflakeSQL
5 Days Ago
Remote
USA
135K-150K Annually
Senior level
135K-150K Annually
Senior level
Healthtech
Build and operate Danaher’s enterprise data platform automation and reliability layer. Responsibilities include Infrastructure-as-Code, CI/CD, self-service developer enablement, data platform provisioning, observability, SLOs, auto-remediation, FinOps, compliance guardrails, and event-driven operations across Snowflake, Azure, Matillion, dbt, and Airflow. The role also develops AI-driven operators and copilots using Azure AI Foundry, Anthropic Claude, and related agent frameworks.
Top Skills: Anthropic ClaudeApache AirflowAutogenAzureAzure Ai FoundryAzure MonitorAzure OpenaiDbtGithub ActionsGitlabLanggraphLog AnalyticsMatillionOpenaiPrompt FlowPulumiPythonSemantic KernelServicenowSnowflakeSnowflake CortexTerraform
6 Days Ago
In-Office or Remote
United States
38K-133K Annually
Senior level
38K-133K Annually
Senior level
Agency • Information Technology
Designs, builds, and operates scalable enterprise data platforms supporting ingestion, processing, storage, governance, analytics, and AI/ML. Responsibilities include developing batch and real-time pipelines, lakehouse and warehouse architectures, orchestration, infrastructure automation, CI/CD, monitoring, security, cost optimization, and observability. The role partners with data, ML, DevOps, security, and application teams, troubleshoots complex platform issues, establishes reusable engineering standards, contributes to architecture strategy, and mentors engineers.
Top Skills: Access ControlAlertingCi/CdCloud Data PlatformsData GovernanceData LakesData OrchestrationData PipelinesData SecurityData WarehousesDevOpsDistributed SystemsEncryptionInfrastructure As CodeLakehousesLoggingMonitoringObservabilityPlatform EngineeringStreamingWorkflow Automation

What you need to know about the NYC Tech Scene

As the undisputed financial capital of the world, New York City is an epicenter of startup funding activity. The city has a thriving fintech scene and is a major player in verticals ranging from AI to biotech, cybersecurity and digital media. It also has universities like NYU, Columbia and Cornell Tech attracting students and researchers from across the globe, providing the ecosystem with a constant influx of world-class talent. And its East Coast location and three international airports make it a perfect spot for European companies establishing a foothold in the United States.

Key Facts About NYC Tech

  • Number of Tech Workers: 549,200; 6% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Capgemini, Bloomberg, IBM, Spotify
  • Key Industries: Artificial intelligence, Fintech
  • Funding Landscape: $25.5 billion in venture capital funding in 2024 (Pitchbook)
  • Notable Investors: Greycroft, Thrive Capital, Union Square Ventures, FirstMark Capital, Tiger Global Management, Tribeca Venture Partners, Insight Partners, Two Sigma Ventures
  • Research Centers and Universities: Columbia University, New York University, Fordham University, CUNY, AI Now Institute, Flatiron Institute, C.N. Yang Institute for Theoretical Physics, NASA Space Radiation Laboratory

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account