Goldman Sachs Logo

Goldman Sachs

Software Engineering - Data, Lakehouse and AI Data Platform Engineer - Vice President - Dallas

Posted 3 Days Ago
Be an Early Applicant
In-Office
Dallas, TX
Senior level
In-Office
Dallas, TX
Senior level
Design, build, test and support batch and streaming data pipelines and curated datasets on a lakehouse/AI data platform. Work across ingestion, transformation, modeling, optimization, data quality and reconciliation; improve platform tooling and standards; collaborate with stakeholders and lead delivery and technical design for workstreams.
The summary above was generated by AI

The Opportunity

Join a team building the data foundations that support the firm’s AI and analytics capabilities. This role sits within the engineering effort to develop a modern Lakehouse and AI data platform that enables reliable, well-governed and high-performing data use across the firm.

At Goldman Sachs, engineering teams are positioned at the center of the business, building scalable systems, solving complex technical problems and turning data into action. In data engineering roles, the emphasis is on designing, building and maintaining large-scale data platforms, delivering production pipelines, improving reliability and quality, and partnering closely with users of the platform.

This is a delivery-focused role for engineers who want to build robust data assets in production, work with modern data technologies, and grow over time within the firm. You will contribute to the data models, pipelines and platform capabilities that underpin analytics, operational decision-making and emerging AI use cases, and may also help extend platform tooling where additional functionality is needed.

Role Summary

As a Data Engineer in the Lakehouse and AI Data Platform team, you will design, build, test and support data pipelines and curated datasets on the firm’s modern data platform. You will work across ingestion, transformation, modelling, optimization and data quality, helping to deliver data products that are reliable, scalable and fit for purpose.  Where there are gaps in platform functionality, you may also contribute to shared tooling or framework components that improve how the platform is used and operated.

The role is suited to engineers who are comfortable writing code, working with SQL and distributed data processing, and solving practical delivery problems in a team environment. More experienced candidates may also contribute to technical design, platform standards and the shaping of delivery approaches across a wider set of use cases.

Key Responsibilities

Pipeline Engineering

  • Build, enhance and support batch and streaming data pipelines on the Lakehouse and AI data platform.
  • Refactor or modernize existing data flows where needed to improve reliability, performance and maintainability.
  • Where needed, build reusable tooling to improve delivery, consistency and operational support.
  • Ensure data pipelines are production-ready, well tested and operationally supportable.

Data Modelling and Curation

  • Develop raw, refined and curated datasets that support analytics, reporting and AI use cases.
  • Apply sound data modelling principles to represent business entities, relationships and historical change accurately.
  • Work with consumers to shape data products that are usable, well documented and aligned to business needs.

Data Quality and Reconciliation

  • Implement controls to validate completeness, accuracy and consistency of data across pipelines and datasets.
  • Use reconciliation approaches to build confidence in production outputs and investigate breaks where they arise.
  • Contribute to clear standards for testing, monitoring and issue resolution.
  • Contribute to practical improvements in testing, monitoring or reconciliation tooling where these strengthen platform reliability and day-to-day delivery.

Delivery and Partnership

  • Work closely with engineers, platform teams and data consumers to deliver agreed outcomes to time and quality expectations.
  • Communicate clearly on progress, risks, dependencies and design choices, including where delivery would benefit from improvements to shared platform tooling.
  • For more senior candidates, take a broader role in technical leadership, task breakdown and support for junior engineers.

Skills and Experience

Required

  • 7-12+ years of experience
  • Bachelor’s or master’s degree in a relevant discipline, or equivalent practical experience, with evidence of strong quantitative skills or data engineering expertise.
  • Strong hands-on programming experience in Python or Java.
  • Good working knowledge of SQL, including troubleshooting, optimization and data analysis.
  • Ability to learn new tools, internal platforms and delivery workflows quickly.
  • Familiarity with software engineering fundamentals, including version control, testing, release discipline and CI/CD practices.

Data Engineering Capability

  • Understanding of temporal data modelling, including the handling of historical state and change over time.
  • Knowledge of schema design, schema evolution and data compatibility considerations.
  • Understanding of partitioning, clustering and other techniques used to improve data performance at scale.
  • Ability to make sensible design choices across normalized and deformalized models, and between natural and surrogate keys.
  • Practical approach to data quality, reconciliation and root-cause analysis.
  • Experience building or supporting production data pipelines in a collaborative engineering environment.
  • Experience working with distributed data processing frameworks such as Apache Spark.
  • Working knowledge of common data formats such as JSONAvro and Parquet.
  • Stronger ownership of technical design across multiple datasets or pipeline domains.
  • Experience guiding implementation standards, code quality and engineering practices within a team.
  • Ability to lead delivery for a workstream, manage dependencies and support less experienced engineers.

Technology Environment

The role will involve working with a modern and evolving data stack. Candidates are not expected to have deep expertise in every tool from day one but should bring relevant experience and the ability to work across comparable technologies.

Examples of technologies in scope include:

  • Data processing and logic: ANSI SQL, Apache Spark, Kafka
  • Data formats: JSON, Avro, Parquet
  • Platforms and storage: Snowflake, Apache Iceberg, Databricks, Hadoop ecosystem technologies, Sybase IQ
  • Engineering and deployment: CI/CD tooling, containerized or Kubernetes-based deployment approaches where relevant

You will also work with internal data management and platform tooling, so a practical and adaptable engineering mindset is important.

What We Are Looking For

We are looking for engineers who can deliver well-structured, reliable solutions in production and who take ownership of the quality of what they build. The role suits candidates who are technically strong, pragmatic and comfortable working in a fast-paced environment where data platforms support important business outcomes.

Stronger candidates will typically demonstrate:

  • sound judgement in technical trade-offs
  • attention to detail in data correctness and testing
  • a clear and structured approach to problem solving
  • willingness to work closely with stakeholders and partner teams
  • an interest in developing long-term expertise within the firm
HQ

Goldman Sachs New York, New York, USA Office

200 West Street, New York, NY, United States, 10282

Goldman Sachs Edison, New Jersey, USA Office

Edison, United States

Goldman Sachs Jersey City, New Jersey, USA Office

Jersey City, United States

Goldman Sachs New York, New York, USA Office

New York, United States

Goldman Sachs Newark, New Jersey, USA Office

Newark, United States

Similar Jobs

8 Minutes Ago
In-Office
Mid level
Mid level
Healthtech • Logistics • Pharmaceutical
Perform detection, investigation, and response to cybersecurity incidents (phishing, malware, ransomware). Analyze logs and forensic data, contain and remediate threats, escalate complex incidents, maintain SOC playbooks, collaborate with threat intel and vulnerability teams, and mentor junior analysts.
Top Skills: CrowdstrikeEdrForensic ToolsIso 27035Mitre Att&CkNistSIEMSplunkWireshark
2 Hours Ago
Hybrid
16-25 Hourly
Junior
16-25 Hourly
Junior
eCommerce • Fashion • Retail • Sales • Wearables • Design
Lead and coach the store sales team to achieve sales targets and deliver excellent customer experiences. Monitor sales data and KPIs, recruit and train staff, provide performance feedback, enforce company policies, support HR/conflict resolution, and implement company initiatives. Serve as a brand ambassador and mentor clienteling and customer-centric strategies. Perform physical tasks including lifting, stocking, and maintaining the sales floor and stockroom.
3 Hours Ago
In-Office
Expert/Leader
Expert/Leader
Artificial Intelligence • Hardware • Information Technology • Machine Learning
The HBM SoC Physical Design Engineer will implement advanced HBM SoC designs, optimize performance and power, and collaborate with various teams to ensure robust physical design integrity.
Top Skills: Cadence InnovusCadence TempusSiemens CalibreSynopsys Icc2Synopsys Primetime

What you need to know about the NYC Tech Scene

As the undisputed financial capital of the world, New York City is an epicenter of startup funding activity. The city has a thriving fintech scene and is a major player in verticals ranging from AI to biotech, cybersecurity and digital media. It also has universities like NYU, Columbia and Cornell Tech attracting students and researchers from across the globe, providing the ecosystem with a constant influx of world-class talent. And its East Coast location and three international airports make it a perfect spot for European companies establishing a foothold in the United States.

Key Facts About NYC Tech

  • Number of Tech Workers: 549,200; 6% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Capgemini, Bloomberg, IBM, Spotify
  • Key Industries: Artificial intelligence, Fintech
  • Funding Landscape: $25.5 billion in venture capital funding in 2024 (Pitchbook)
  • Notable Investors: Greycroft, Thrive Capital, Union Square Ventures, FirstMark Capital, Tiger Global Management, Tribeca Venture Partners, Insight Partners, Two Sigma Ventures
  • Research Centers and Universities: Columbia University, New York University, Fordham University, CUNY, AI Now Institute, Flatiron Institute, C.N. Yang Institute for Theoretical Physics, NASA Space Radiation Laboratory

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account