Goldman Sachs Logo

Goldman Sachs

Vice President - Site Reliability Engineering (SRE) – The Core Engineering

Posted 21 Hours Ago
Be an Early Applicant
In-Office
New York, NY, USA
Expert/Leader
In-Office
New York, NY, USA
Expert/Leader
Lead Site Reliability Engineering initiatives for critical, large-scale financial platforms. Responsibilities include defining SLOs, SLIs, and error budgets; designing resilient distributed systems; developing automation and self-service tooling; improving production readiness through testing and capacity planning; leading complex incident response and blameless post-mortems; and promoting sustainable on-call practices across engineering teams.
The summary above was generated by AI
Vice President - Site Reliability Engineering (SRE) – The Core Engineering
WHAT WE DO

Site Reliability Engineering at Goldman Sachs sits at the intersection of software engineering, systems design, and production excellence. In this VP role, you will help engineer highly reliable, observable, and resilient platforms that support critical business services at scale. You will collaborate with multiple engineering teams to continually improve our production system architecture, facilitate fast delivery of new services, and reduce downtime.

This role is for software engineers who enjoy solving complex distributed system problems, building tools and platforms that make teams more effective, and championing SRE principles (such as SLOs, error budgets, and blameless post-mortems) across a large engineering organization.

Key Responsibilities
  • Partner with engineering leadership to establish service level objectives (SLOs), service level indicators (SLIs), and error budgets.
  • Collaborate with product developers to architect highly available, fault-tolerant, and self-healing systems. Conduct architectural reviews and introduce patterns like circuit breakers, graceful degradation, and rate limiting.
  • Reduce operational toil by building automation, tooling, and self-service capabilities that remove repetitive manual work.

  • Improve production readiness through load testing, performance tuning, capacity forecasting, and reliability reviews.

  • Lead the response to complex, multi-system production incidents. Facilitate blameless post-mortems to identify root causes and drive long-term preventative actions.

  • Promote sustainable operations by helping design healthy on-call models, clear escalation paths, and balanced pager responsibilities.


WHAT WE ARE LOOKING FORCore Technical Skills
  • Strong proficiency in at least one major programming language (e.g., Java, Python, or Node.js) with a focus on writing clean, maintainable code for tooling and automation.
  • Hands-on experience with Infrastructure as Code (IaC) frameworks such as Terraform, Ansible, or CloudFormation.

  • Deep understanding of containerization and orchestration technologies, specifically Docker and Kubernetes (K8s), including service meshes and ingress controllers.

  • Advanced experience with major cloud providers (AWS, GCP, or Azure), specifically building and operating highly resilient cloud-native architectures.

  • Proficiency with Observability stacks, including distributed tracing, logging, and metrics (e.g., Prometheus, Grafana, Splunk, Datadog, OpenTelemetry, ELK, or CloudWatch)
  • Experience with automated testing and SDLC concepts, developing applications in a Linux environment, and sound knowledge of algorithms, data structures and software design.

  • Knowledge of networking protocols and load balancing strategies in a distributed systems environment.

Core Competencies & Soft Skills
  • Ability to analyze complex, distributed systems holistically and understand how individual components interact under load.

  • Strong interpersonal skills to collaborate with product developers, influence architectural decisions, prioritize toil reduction, and drive SRE adoption without direct authority.

  • Ability to translate complex technical issues into clear, actionable insights for both technical and non-technical stakeholders.
  • Highly motivated, pro-active and capable of multi-tasking under pressure in a fast-paced environment without compromising quality.

  • Commitment to fostering a blameless culture where failures are treated as opportunities to learn and improve systems.

  • Interest in financial markets and technology.

Preferred Qualifications
  • Bachelor’s degree in Computer Science, System Engineering, or a related technical field that involves programming.

  • 7 to 10 years of experience 


ABOUT GOLDMAN SACHS

The Goldman Sachs Group, Inc. is a leading global investment banking, securities and investment management firm that provides a wide range of financial services to a substantial and diversified client base that includes corporations, financial institutions, governments and individuals. Founded in 1869, the firm is headquartered in New York and maintains offices in all major financial centers around the world.


HQ

Goldman Sachs New York, New York, USA Office

200 West Street, New York, NY, United States, 10282

Goldman Sachs Edison, New Jersey, USA Office

Edison, United States

Goldman Sachs Jersey City, New Jersey, USA Office

Jersey City, United States

Goldman Sachs New York, New York, USA Office

New York, United States

Goldman Sachs Newark, New Jersey, USA Office

Newark, United States

Similar Jobs

3 Minutes Ago
In-Office
36K-82K Annually
Junior
36K-82K Annually
Junior
Healthtech • Logistics • Pharmaceutical
Perform routine IT operations tasks to deploy, monitor, and maintain enterprise systems. Follow documented troubleshooting steps, resolve standard alerts, escalate unclear issues, and assist senior admins with lifecycle tasks. Communicate within IT operations, adhere to governance, and support system patching and log analysis under supervision.
Top Skills: AWSLinuxAzureVirtualizationWindows
7 Minutes Ago
Easy Apply
Hybrid
New York, NY, USA
Easy Apply
104K-155K Annually
Senior level
104K-155K Annually
Senior level
Fintech • HR Tech
Support health-insurance sales by building reports, forecasts, automations, and AI prototypes to improve pipeline health, rep productivity, and territory/quota planning. Partner cross-functionally to maintain data integrity, troubleshoot CRM issues, administer account and territory assignments, and drive process improvements and adoption.
Top Skills: Ai ToolsExcelGoogle SheetsLookerSalesforceSQLTableau
8 Minutes Ago
In-Office or Remote
New York, NY, USA
50K-72K Annually
Mid level
50K-72K Annually
Mid level
Artificial Intelligence • Cloud • Fintech • Information Technology • Insurance • Financial Services • Big Data Analytics
Manages an assigned long-term disability claims caseload by evaluating eligibility, reviewing medical and contractual information, calculating benefits, and making timely claim decisions. Develops strategic case plans, communicates with customers, employers, physicians, attorneys, and vocational counselors, and facilitates return-to-work opportunities. Maintains claim documentation, meets service and performance standards, follows compliance procedures, and handles complex customer interactions while coordinating with internal clinical and business partners.
Top Skills: ExcelMS OfficeMicrosoft OutlookMicrosoft Word

What you need to know about the NYC Tech Scene

As the undisputed financial capital of the world, New York City is an epicenter of startup funding activity. The city has a thriving fintech scene and is a major player in verticals ranging from AI to biotech, cybersecurity and digital media. It also has universities like NYU, Columbia and Cornell Tech attracting students and researchers from across the globe, providing the ecosystem with a constant influx of world-class talent. And its East Coast location and three international airports make it a perfect spot for European companies establishing a foothold in the United States.

Key Facts About NYC Tech

  • Number of Tech Workers: 549,200; 6% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Capgemini, Bloomberg, IBM, Spotify
  • Key Industries: Artificial intelligence, Fintech
  • Funding Landscape: $25.5 billion in venture capital funding in 2024 (Pitchbook)
  • Notable Investors: Greycroft, Thrive Capital, Union Square Ventures, FirstMark Capital, Tiger Global Management, Tribeca Venture Partners, Insight Partners, Two Sigma Ventures
  • Research Centers and Universities: Columbia University, New York University, Fordham University, CUNY, AI Now Institute, Flatiron Institute, C.N. Yang Institute for Theoretical Physics, NASA Space Radiation Laboratory

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account