Wells Fargo Logo

Wells Fargo

Senior Lead Platform Reliability Engineer

Reposted 7 Hours Ago
Be an Early Applicant
Hybrid
Iselin, NJ, USA
159K-305K Annually
Senior level
Hybrid
Iselin, NJ, USA
159K-305K Annually
Senior level
Lead reliability engineering for a primary infrastructure domain (Network, Middleware, Database, or Storage). Apply SRE practices to improve availability, resiliency, observability, capacity planning, automation, and incident response. Drive root-cause analysis, performance tuning, automation to remove toil, define observability standards, and mentor engineers to improve platform stability at scale.
The summary above was generated by AI
Wells Fargo is seeking a Senior Lead Platform Reliability Engineer to join the CTO Platform organization. This role is designed for highly experienced infrastructure engineers who possess deep technical expertise in one core platform discipline (Network, Middleware, Database, or Storage) and have demonstrated experience collaborating across at least one additional infrastructure domains (Enterprise Tools, Cloud, Observability Tools). The expectation is that this engineer will elevate themselves in looking for trends and patterns that are not limited to these streams and investigate systemic issues that span multiple streams or need a deeper troubleshooting.
As part of our Platform Reliability Engineering (PRE) team, you will apply modern Site Reliability Engineering (SRE) practices to improve the availability, resiliency, observability, scalability, and operational excellence of critical enterprise platforms. You will leverage your domain expertise to identify systemic issues, drive automation, and deliver engineering solutions that strengthen platform stability at scale.
In This Role You Will
  • Serve as the reliability engineering expert for your primary domain (Network, Middleware, Database, or Storage) while partnering across adjacent technology disciplines
  • Lead the investigation and resolution of complex production incidents, identifying root causes and implementing long-term corrective actions
  • Apply SRE principles including service level indicators (SLIs), service level objectives (SLOs), error budgets, and reliability engineering practices to improve platform health
  • Lead capacity analysis, forecasting, and utilization reviews to identify future scaling risks and prevent service degradation before customer impact occurs
  • Perform deep performance analysis across infrastructure layers, identifying bottlenecks, contention points, latency drivers, and resource inefficiencies
  • Identify and remediate configuration drift, operational debt, and platform hygiene issues that impact long-term reliability
  • Drive proactive reliability improvements through observability, automation, performance optimization, and resiliency engineering
  • Design and implement automation solutions that eliminate operational toil, reduce manual intervention, and improve recovery capabilities
  • Define and enhance enterprise observability standards through metrics, logging, tracing, alerting, and service health monitoring
  • Partner closely with engineering, infrastructure, application, cloud, and operations teams to improve platform performance and availability
  • Lead blameless post-incident reviews and convert recurring operational issues into measurable engineering improvements
  • Identify reliability risks and communicate technical recommendations to engineering leaders and senior stakeholders
  • Mentor engineers and technical teams on reliability engineering, operational excellence, automation, and platform best practices
Required Qualifications
  • 7+ years of Systems Engineering, Technology Architecture experience, or equivalent demonstrated through one or a combination of the following: work experience, training, military experience, education
  • 5+ years supporting and engineering enterprise-scale production environments
  • 5+ years of experience with hands-on expertise in one of the following technology domains:
    • Network Engineering (routing, switching, load balancing, DNS, network observability, performance analysis)
    • Middleware Engineering (WebSphere, Tomcat, JBoss, Kafka, MQ, application platforms, integration technologies)
    • Database Engineering (Oracle, SQL Server, PostgreSQL, MongoDB, database performance, replication, HA/DR)
    • Storage Engineering (SAN/NAS technologies, storage virtualization, backup/recovery, performance and capacity management)
Desired Qualifications
  • Strong experience applying SRE principles, including SLI/SLO development, error budgets, incident analysis, and reliability measurement
  • Experience supporting highly available, mission-critical production environments
  • Proven success troubleshooting complex issues spanning multiple technology domains in large-scale distributed environments
  • Experience with capacity planning, resiliency engineering, fault tolerance, disaster recovery, and performance optimization
  • Hands-on experience with observability and monitoring platforms such as Grafana, Splunk, Prometheus, AppDynamics, Cribl, ThousandEyes, Dynatrace, or similar technologies
  • Experience building dashboards, alerts, service health indicators, and operational reporting
  • Strong automation and scripting experience using Python, Bash, PowerShell, or similar technologies
  • Experience developing operational tooling, API integrations, self-healing capabilities, and automated remediation solutions
  • Familiarity with Git-based development practices, CI/CD pipelines, infrastructure automation, and Infrastructure as Code tools such as Ansible or Terraform
  • Experience diagnosing and resolving issues that span multiple infrastructure layers
  • Ability to influence technical direction across infrastructure and engineering organizations
  • Experience leading major incident reviews and driving sustainable operational improvements
  • Demonstrated success mentoring engineers and promoting reliability engineering best practices
  • Strong communication skills with the ability to translate technical concepts into business-focused outcomes
Job Expectations:
  • This position offers a hybrid schedule
  • This position does not offer Visa sponsorship
Pay Range
Reflected is the base pay range offered for this position. Pay may vary depending on factors including but not limited to demonstrated examples of prior performance, skills, experience, or work location. Employees may also be eligible for incentive opportunities.
$159,000.00 - $305,000.00
Benefits
Wells Fargo provides eligible employees with a comprehensive set of benefits, many of which are listed below. Visit Benefits - Wells Fargo Jobs for an overview of the following benefit plans and programs offered to employees.
  • Health benefits
  • 401(k) Plan
  • Paid time off
  • Disability benefits
  • Life insurance, critical illness insurance, and accident insurance
  • Parental leave
  • Critical caregiving leave
  • Discounts and savings
  • Commuter benefits
  • Tuition reimbursement
  • Scholarships for dependent children
  • Adoption reimbursement
Posting End Date:
30 Aug 2026
* Job posting may come down early due to volume of applicants.
We Value Equal Opportunity
Wells Fargo is an equal opportunity employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, disability, status as a protected veteran, or any other legally protected characteristic.
Employees support our focus on building strong customer relationships balanced with a strong risk mitigating and compliance-driven culture which firmly establishes those disciplines as critical to the success of our customers and company. They are accountable for execution of all applicable risk programs (Credit, Market, Financial Crimes, Operational, Regulatory Compliance), which includes effectively following and adhering to applicable Wells Fargo policies and procedures, appropriately fulfilling risk and compliance obligations, timely and effective escalation and remediation of issues, and making sound risk decisions. There is emphasis on proactive monitoring, governance, risk identification and escalation, as well as making sound risk decisions commensurate with the business unit's risk appetite and all risk and compliance program requirements.
Applicants with Disabilities
To request a medical accommodation during the application or interview process, visit Disability Inclusion at Wells Fargo .
Drug and Alcohol Policy
Wells Fargo maintains a drug free workplace. Please see our Drug and Alcohol Policy to learn more.
Wells Fargo Recruitment and Hiring Requirements:
a. Third-Party recordings are prohibited unless authorized by Wells Fargo.
b. Wells Fargo requires you to directly represent your own experiences during the recruiting and hiring process.
#DNP-IND
#BI-Hybrid

Wells Fargo New York, New York, USA Office

150 E 42nd Street, New York, NY, United States, 10017

Wells Fargo New York, New York, USA Office

500 West 33rd Street Manhattan, New York, NY, United States, 10001

Similar Jobs at Wells Fargo

7 Hours Ago
Hybrid
Iselin, NJ, USA
87K-168K Annually
Junior
87K-168K Annually
Junior
Fintech • Financial Services
Supports cybersecurity Identity and Access Management through data engineering, analytics, and machine learning. Responsibilities include building and operating data pipelines, models, dashboards, and API integrations; developing, testing, deploying, and monitoring ML-enabled solutions; evaluating model performance and drift; ensuring compliance with security, privacy, regulatory, and ethical requirements; using AI-assisted development tools; and communicating technical recommendations to IAM stakeholders and business partners.
Top Skills: AlteryxAPIsClaude CodeCursorDevinGitGithub CopilotGoogle Cloud PlatformLlmsNumpyPandasPower BIPythonPyTorchRagScikit-LearnSQLTableauTensorFlowVertex Ai
7 Hours Ago
Hybrid
Iselin, NJ, USA
159K-305K Annually
Senior level
159K-305K Annually
Senior level
Fintech • Financial Services
Leads technology control modernization by redesigning controls, aligning definitions with RCSA, and embedding governance, security, and automation across the technology lifecycle. Advises senior leaders on complex and emerging risks, evaluates control environments, monitors effectiveness and compliance data, supports remediation, reports risk outcomes, and mentors control management teams. Collaborates across business lines to implement risk mitigation strategies, address regulatory requirements, and improve automated configuration monitoring.
Top Skills: AIAutomationConfiguration MonitoringIt Systems SecurityRcsa
7 Hours Ago
Hybrid
Iselin, NJ, USA
159K-305K Annually
Senior level
159K-305K Annually
Senior level
Fintech • Financial Services
Lead strategy, architecture, and roadmap for enterprise application and web server platforms (Tomcat, WebSphere, NGINX). Drive automation, IaC, CI/CD, and self-service platform engineering; advise leadership and collaborate with architecture, security, cloud, and application teams to govern middleware lifecycle and standards.
Top Skills: AnsibleApache Http ServerApache TomcatArtifactoryAzure DevopsBashGithub ActionsGitlabIbm Websphere LibertyJenkinsLinuxNginxPowershellPythonTerraform

What you need to know about the NYC Tech Scene

As the undisputed financial capital of the world, New York City is an epicenter of startup funding activity. The city has a thriving fintech scene and is a major player in verticals ranging from AI to biotech, cybersecurity and digital media. It also has universities like NYU, Columbia and Cornell Tech attracting students and researchers from across the globe, providing the ecosystem with a constant influx of world-class talent. And its East Coast location and three international airports make it a perfect spot for European companies establishing a foothold in the United States.

Key Facts About NYC Tech

  • Number of Tech Workers: 549,200; 6% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Capgemini, Bloomberg, IBM, Spotify
  • Key Industries: Artificial intelligence, Fintech
  • Funding Landscape: $25.5 billion in venture capital funding in 2024 (Pitchbook)
  • Notable Investors: Greycroft, Thrive Capital, Union Square Ventures, FirstMark Capital, Tiger Global Management, Tribeca Venture Partners, Insight Partners, Two Sigma Ventures
  • Research Centers and Universities: Columbia University, New York University, Fordham University, CUNY, AI Now Institute, Flatiron Institute, C.N. Yang Institute for Theoretical Physics, NASA Space Radiation Laboratory

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account