Goldman Sachs Logo

Goldman Sachs

Vice President - AI Safety Platform Engineering

Posted 21 Days Ago
Be an Early Applicant
In-Office
New York, NY, USA
130K-250K Annually
Senior level
In-Office
New York, NY, USA
130K-250K Annually
Senior level
Lead enterprise AI safety engineering by building agent evaluation frameworks, real-time LLM guardrails, governance controls, telemetry, audit logging, and human-approval workflows. Recruit and mentor engineering teams, own the AI safety infrastructure roadmap, partner with core AI platform and risk teams, and translate financial-services model risk and regulatory requirements into automated safeguards. The role requires deep software engineering, agentic AI, evaluation, observability, and platform leadership experience.
The summary above was generated by AI

Role Overview 

We are seeking aVice President – AI Safety Platformsto build and lead our enterprise AI safety engineering initiatives. As generative AI in financial services evolves from simple prompt-response workflows to autonomous agentic systems that execute multi-step plans, call APIs, and interact directly with internal systems, establishing robust safety mechanisms and standardized evaluation protocols is essential. 

In this role, you will recruit and lead dedicated engineering pods focused on developinga unified company-wide agentic evaluation framework, real-time LLM guardrail services, and automated governance controls. As the senior technical authority for AI safety, you will collaborate closely with core AI platform teams, risk control functions, and business units to drive necessary enhancements to the core AI platform (such as telemetry hooks, API capabilities, execution sandboxes, and data logging infrastructure) to ensure all enterprise AI deployments operate safely, verifiably, and in compliance with institutional standards. 

 

Key Responsibilities 

1. Unified Agentic Evaluation Framework 

  • Company-Wide Architecture:Design, build, and deploya single, company-wide agentic evaluation frameworkthat standardizes how teams across all business lines benchmark, test, and measure AI agent performance prior to production deployment. 
  • Trajectory & Multi-Step Reasoning Assessment:Implement evaluation methodologies that score autonomous planning quality, tool-calling precision, multi-turn state retention, trajectory efficiency, and error-recovery behaviors. 
  • Continuous Monitoring & Production Drift:Integrate automated evaluation pipelines into runtime environments to continuously audit agent execution traces, detecting reasoning drift, tool failure modes, and unexpected trajectory shifts in production. 
  • Domain-Specific Benchmarking:Establish standardized test suites and synthetic evaluation benchmarks tailored to complex financial workflows, such as automated research, risk assessment, and operational task execution. 

2. LLM Guardrails Infrastructure & Real-Time Controls 

  • Low-Latency Guardrail Engine:Architect and scale enterprise guardrail microservices that inspect prompt inputs, retrieved context, and model outputs in real time to prevent data leakage, policy violations, and unvalidated execution. 
  • Tool-Use & Action Control:Implement runtime policy gateways that inspect and authorize tool calls before execution, ensuring agents operate within authorized data boundaries and action scopes. 
  • Human-in-the-Loop (HITL) Triggers:Build configurable escalation workflows and approval gates that automatically pause execution for high-risk operations (e.g., money movement, client record modifications, or external communications) until human authorization is granted. 

3. Core AI Platform Enhancements & Governance Integration 

  • Drive Platform Enhancements:Partner directly with the core AI Platform team to drive the implementation of safety APIs, telemetry hooks, developer SDKs, and MLOps/LLMOps pipeline integrations. 
  • Auditability & Execution Telemetry:Define and enforce technical standards for immutable audit logging, execution tracing (e.g., OpenTelemetry standards), and principal identity propagation across all agentic workflows. 
  • Regulatory & Model Risk Alignment:Translate model risk management standards (e.g., SR 11-7 / SR 26-2 guidance, FINRA supervision requirements) into automated engineering safeguards and policy checks. 

4. Engineering Leadership & Strategic Oversight 

  • Team Building & Mentorship:Hire, develop, and mentor high-performing engineering teams specializing in applied machine learning, AI safety, and enterprise platform engineering. 
  • Strategic Roadmap:Own the technical roadmap for enterprise AI safety infrastructure, setting clear milestones for evaluation framework adoption, runtime latency optimization, and governance automation. 
  • Stakeholder Collaboration:Articulate technical risk profiles, evaluation metrics, and safety architecture to risk committees, model validation teams, and executive leadership. 

 

Key Qualifications 

Basic Qualifications 

  • Role Level:Vice President experience (or equivalent senior engineering leadership) in financial services or large-scale enterprise software environments. 
  • Education:Bachelor’s or Master’s degree in Computer Science, Artificial Intelligence, Systems Engineering, or a related quantitative field. 
  • Engineering Leadership:4+ years leading applied ML or software engineering teams in building platform infrastructure or microservices. 
  • Software Engineering Depth:8+ years of hands-on software development experience (Python, Go, Java, or C++) building microservices, high-throughput APIs, or enterprise platform services. 
  • AI & Agentic Expertise:Technical fluency with Large Language Models (LLMs), RAG systems, function calling / tool integration, and agentic execution paradigms (e.g., LangChain, AutoGen, CrewAI, MCP server architectures). 

Preferred Experience & Technical Skills 

  • Agentic Evaluation:Direct experience building agent evaluation frameworks and metrics (e.g., LLM-as-a-Judge, G-Eval, trajectory trace evaluation, task completion scoring). 
  • Guardrail Frameworks:Hands-on experience integrating low-latency guardrail tools and runtime filters (e.g., NeMo Guardrails, Guardrails AI, Llama Guard). 
  • AI Observability & Tracing:Experience with LLM and agent tracing tools (e.g., LangSmith, OpenTelemetry, Phoenix, MLflow) and structured audit logging infrastructure. 
  • Platform Engineering Alignment:Proven ability to partner across teams and drive key governance capabilities into core shared platforms. 

Salary Range
The expected base salary for this New York, NY, United States-based position is $130000-$250000. In addition, you may be eligible for a discretionary bonus if you are an active employee as of fiscal year-end.

Benefits
Goldman Sachs is committed to providing our people with valuable and competitive benefits and wellness offerings, as it is a core part of providing a strong overall employee experience. A summary of these offerings, which are generally available to active, non-temporary, full-time and part-time US employees who work at least 20 hours per week, can be found here.

HQ

Goldman Sachs New York, New York, USA Office

200 West Street, New York, NY, United States, 10282

Goldman Sachs Edison, New Jersey, USA Office

Edison, United States

Goldman Sachs Jersey City, New Jersey, USA Office

Jersey City, United States

Goldman Sachs New York, New York, USA Office

New York, United States

Goldman Sachs Newark, New Jersey, USA Office

Newark, United States

Similar Jobs

2 Minutes Ago
In-Office or Remote
New York, NY, USA
87K-130K Hourly
Expert/Leader
87K-130K Hourly
Expert/Leader
Blockchain • eCommerce • Fintech • Payments • Software • Financial Services • Cryptocurrency
Provides strategic and administrative partnership to senior executives, managing calendars, travel, expenses, operating rhythms, meetings, information flow, and follow-through. Improves workflows through AI, automation, documentation, reporting, and scalable systems. Builds cross-functional relationships, supports executive programs and planning processes, mentors peers, and helps strengthen centralized Executive Operations practices. Principal-level candidates also architect support models, lead process improvements, and coach executive support professionals.
Top Skills: AIAutomation ToolsGoogle CalendarGoogle DocsGoogle FormsGoogle SheetsGoogle SlidesGoogle WorkspaceWorkflow Tools
2 Minutes Ago
In-Office or Remote
New York, NY, USA
264K-395K Annually
Expert/Leader
264K-395K Annually
Expert/Leader
Blockchain • eCommerce • Fintech • Payments • Software • Financial Services • Cryptocurrency
Build and own foundational Android platform infrastructure, including architecture, shared libraries, developer tooling, observability, performance, and experimentation systems. Lead AI-native engineering practices, evaluate AI development tools, and drive cross-organizational technical initiatives. Partner across engineering, product, and design; establish scalable standards; identify systemic risks; communicate architectural tradeoffs; and mentor Android engineers across the organization.
Top Skills: Ai-Assisted Development ToolsAndroidArtifactoryBuildkiteClaude CodeCodexCustom Gradle PluginsGooseGradleJetpack ComposeKochikuKotlinKotlin MultiplatformMoleculePaparazziSqldelightStructured ConcurrencyWireWorkmanager
23 Minutes Ago
Remote or Hybrid
United States
17-25 Hourly
Junior
17-25 Hourly
Junior
Artificial Intelligence • Automotive • Greentech • Information Technology • Machine Learning • Software • Cybersecurity
Provides remote technical customer support for Dealertrack and Cox Automotive products through phone, email, and chat. Resolves routine application and product issues, documents cases in CRM, trains customers on product usage, coordinates with internal teams, and follows up through resolution. Requires strong troubleshooting, communication, customer service, and documentation skills, with flexibility for variable shifts, Saturdays, and overtime.
Top Skills: Cox AutomotiveCRMDealertrackGenesys PurecloudSalesforce

What you need to know about the NYC Tech Scene

As the undisputed financial capital of the world, New York City is an epicenter of startup funding activity. The city has a thriving fintech scene and is a major player in verticals ranging from AI to biotech, cybersecurity and digital media. It also has universities like NYU, Columbia and Cornell Tech attracting students and researchers from across the globe, providing the ecosystem with a constant influx of world-class talent. And its East Coast location and three international airports make it a perfect spot for European companies establishing a foothold in the United States.

Key Facts About NYC Tech

  • Number of Tech Workers: 549,200; 6% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Capgemini, Bloomberg, IBM, Spotify
  • Key Industries: Artificial intelligence, Fintech
  • Funding Landscape: $25.5 billion in venture capital funding in 2024 (Pitchbook)
  • Notable Investors: Greycroft, Thrive Capital, Union Square Ventures, FirstMark Capital, Tiger Global Management, Tribeca Venture Partners, Insight Partners, Two Sigma Ventures
  • Research Centers and Universities: Columbia University, New York University, Fordham University, CUNY, AI Now Institute, Flatiron Institute, C.N. Yang Institute for Theoretical Physics, NASA Space Radiation Laboratory

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account