Goldman Sachs Logo

Goldman Sachs

Vice President - Site Reliability Engineering (SRE) – The Core Engineering

Posted 21 Hours Ago
Be an Early Applicant
In-Office
New York, NY, USA
Expert/Leader
In-Office
New York, NY, USA
Expert/Leader
Lead Site Reliability Engineering initiatives for critical, large-scale financial platforms. Responsibilities include defining SLOs, SLIs, and error budgets; designing resilient distributed systems; developing automation and self-service tooling; improving production readiness through testing and capacity planning; leading complex incident response and blameless post-mortems; and promoting sustainable on-call practices across engineering teams.
The summary above was generated by AI
Vice President - Site Reliability Engineering (SRE) – The Core Engineering
WHAT WE DO

Site Reliability Engineering at Goldman Sachs sits at the intersection of software engineering, systems design, and production excellence. In this VP role, you will help engineer highly reliable, observable, and resilient platforms that support critical business services at scale. You will collaborate with multiple engineering teams to continually improve our production system architecture, facilitate fast delivery of new services, and reduce downtime.

This role is for software engineers who enjoy solving complex distributed system problems, building tools and platforms that make teams more effective, and championing SRE principles (such as SLOs, error budgets, and blameless post-mortems) across a large engineering organization.

Key Responsibilities
  • Partner with engineering leadership to establish service level objectives (SLOs), service level indicators (SLIs), and error budgets.
  • Collaborate with product developers to architect highly available, fault-tolerant, and self-healing systems. Conduct architectural reviews and introduce patterns like circuit breakers, graceful degradation, and rate limiting.
  • Reduce operational toil by building automation, tooling, and self-service capabilities that remove repetitive manual work.

  • Improve production readiness through load testing, performance tuning, capacity forecasting, and reliability reviews.

  • Lead the response to complex, multi-system production incidents. Facilitate blameless post-mortems to identify root causes and drive long-term preventative actions.

  • Promote sustainable operations by helping design healthy on-call models, clear escalation paths, and balanced pager responsibilities.


WHAT WE ARE LOOKING FORCore Technical Skills
  • Strong proficiency in at least one major programming language (e.g., Java, Python, or Node.js) with a focus on writing clean, maintainable code for tooling and automation.
  • Hands-on experience with Infrastructure as Code (IaC) frameworks such as Terraform, Ansible, or CloudFormation.

  • Deep understanding of containerization and orchestration technologies, specifically Docker and Kubernetes (K8s), including service meshes and ingress controllers.

  • Advanced experience with major cloud providers (AWS, GCP, or Azure), specifically building and operating highly resilient cloud-native architectures.

  • Proficiency with Observability stacks, including distributed tracing, logging, and metrics (e.g., Prometheus, Grafana, Splunk, Datadog, OpenTelemetry, ELK, or CloudWatch)
  • Experience with automated testing and SDLC concepts, developing applications in a Linux environment, and sound knowledge of algorithms, data structures and software design.

  • Knowledge of networking protocols and load balancing strategies in a distributed systems environment.

Core Competencies & Soft Skills
  • Ability to analyze complex, distributed systems holistically and understand how individual components interact under load.

  • Strong interpersonal skills to collaborate with product developers, influence architectural decisions, prioritize toil reduction, and drive SRE adoption without direct authority.

  • Ability to translate complex technical issues into clear, actionable insights for both technical and non-technical stakeholders.
  • Highly motivated, pro-active and capable of multi-tasking under pressure in a fast-paced environment without compromising quality.

  • Commitment to fostering a blameless culture where failures are treated as opportunities to learn and improve systems.

  • Interest in financial markets and technology.

Preferred Qualifications
  • Bachelor’s degree in Computer Science, System Engineering, or a related technical field that involves programming.

  • 7 to 10 years of experience 


ABOUT GOLDMAN SACHS

The Goldman Sachs Group, Inc. is a leading global investment banking, securities and investment management firm that provides a wide range of financial services to a substantial and diversified client base that includes corporations, financial institutions, governments and individuals. Founded in 1869, the firm is headquartered in New York and maintains offices in all major financial centers around the world.


HQ

Goldman Sachs New York, New York, USA Office

200 West Street, New York, NY, United States, 10282

Goldman Sachs Edison, New Jersey, USA Office

Edison, United States

Goldman Sachs Jersey City, New Jersey, USA Office

Jersey City, United States

Goldman Sachs New York, New York, USA Office

New York, United States

Goldman Sachs Newark, New Jersey, USA Office

Newark, United States

Similar Jobs

58 Seconds Ago
Hybrid
New York, NY, USA
80K-110K Annually
Senior level
80K-110K Annually
Senior level
AdTech • Big Data • Digital Media • Software
Manages technical escalations for Magnite’s DV+ programmatic advertising platform. Queries large datasets with SQL, troubleshoots issues across systems, APIs, reporting pipelines, and integrations, and applies AI tools to automate workflows and improve efficiency. The role coordinates with Product, Engineering, Account Management, Technical Operations, and external partners while documenting solutions, diagnosing root causes, and driving ambiguous issues to resolution.
Top Skills: Ai ToolsAPIsExcelSQL
A Minute Ago
Hybrid
New York, NY, USA
130K-145K Annually
Junior
130K-145K Annually
Junior
AdTech • Big Data • Digital Media • Software
Develop large-scale distributed systems and high-performance, low-latency software for ad traffic optimization. Apply mathematical algorithms and large-scale data analysis to maximize throughput, improve traffic prioritization, and reduce infrastructure costs. Own the full software development lifecycle, participate in production support and on-call rotations, and collaborate on technical tradeoffs. The role uses Rust, Scala, Spark, Kafka, Kubernetes, and cloud deployment tooling in a hybrid Boston or New York City environment.
Top Skills: AirflowArgo WorkflowsCncfHelmKafkaKubernetesPrometheusRustScalaSpark Structured StreamingTerraform
3 Minutes Ago
In-Office
Brooklyn, NY, USA
20-34 Hourly
Senior level
20-34 Hourly
Senior level
Other • Utilities
Provide best-in-class in-store customer service and sales, resolve customer issues in one interaction using digital tools, complete ongoing training, act as a subject-matter expert, and collaborate with teammates to drive store results. Must be bilingual (English + Spanish or Chinese).

What you need to know about the NYC Tech Scene

As the undisputed financial capital of the world, New York City is an epicenter of startup funding activity. The city has a thriving fintech scene and is a major player in verticals ranging from AI to biotech, cybersecurity and digital media. It also has universities like NYU, Columbia and Cornell Tech attracting students and researchers from across the globe, providing the ecosystem with a constant influx of world-class talent. And its East Coast location and three international airports make it a perfect spot for European companies establishing a foothold in the United States.

Key Facts About NYC Tech

  • Number of Tech Workers: 549,200; 6% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Capgemini, Bloomberg, IBM, Spotify
  • Key Industries: Artificial intelligence, Fintech
  • Funding Landscape: $25.5 billion in venture capital funding in 2024 (Pitchbook)
  • Notable Investors: Greycroft, Thrive Capital, Union Square Ventures, FirstMark Capital, Tiger Global Management, Tribeca Venture Partners, Insight Partners, Two Sigma Ventures
  • Research Centers and Universities: Columbia University, New York University, Fordham University, CUNY, AI Now Institute, Flatiron Institute, C.N. Yang Institute for Theoretical Physics, NASA Space Radiation Laboratory

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account