Lead Platform Reliability Engineer
Overview
You will own reliability for critical enterprise platforms at Wells Fargo, serving as the SRE expert for Network, Middleware, Database, or Storage. You will team with the CTO Platform organization to drive systemic improvements across Cloud, Observability, and Enterprise Tools. Your focus: apply Site Reliability Engineering practices to boost availability, resiliency, and scalability. You will lead incident investigations, capacity planning, and automation to eliminate toil and prevent degradation. This role stands apart by blending deep domain expertise with cross-domain troubleshooting, influencing technology direction at scale.
What You'll Do10
- 1Serve as the reliability engineering expert for your primary domain, partnering across adjacent disciplines like Cloud and Observability.
- 2Lead complex production incident investigations, identifying root causes and implementing long-term corrective actions.
- 3Apply SRE principles including SLIs, SLOs, and error budgets to improve platform health.
- 4Lead capacity analysis and forecasting to prevent service degradation before impact.
- 5Perform deep performance analysis across infrastructure layers, identifying bottlenecks and latency drivers.
- 6Remediate configuration drift and operational debt that impact long-term reliability.
- 7Drive proactive reliability improvements through observability, automation, and resiliency engineering.
- 8Design automation solutions to eliminate toil and improve recovery capabilities, using Python, Bash, or PowerShell.
- 9Define enterprise observability standards with Grafana, Splunk, Prometheus, and similar tools.
- 10Mentor engineers on reliability best practices and lead blameless post-incident reviews.
Requirements10
- 15+ years in Systems Engineering, Infrastructure Engineering, or Platform Engineering.
- 25+ years supporting and engineering enterprise-scale production environments.
- 35+ years hands-on expertise in one domain: Network Engineering, Middleware Engineering, Database Engineering, or Storage Engineering.
- 4Strong experience applying SRE principles, including SLI/SLO development and error budgets.
- 5Experience with observability platforms: Grafana, Splunk, Prometheus, AppDynamics, Dynatrace, or similar.
- 6Hands-on automation and scripting using Python, Bash, or PowerShell.
- 7Familiarity with CI/CD pipelines and Infrastructure as Code tools like Ansible or Terraform.
- 8Proven success troubleshooting complex issues across multiple technology domains.
- 9Experience with capacity planning, resiliency engineering, and disaster recovery.
- 10Strong communication skills to translate technical concepts into business outcomes.
Salary Insight
$119 - $224k per year
Location
Required Skills
Similar open positions
Explore active roles that match your skills and interests.
Wells Fargo
VerifiedSenior Lead Systems Operations Engineer, SRE & Reliability
You will own reliability engineering for Wells Fargo's Commercial and Corporate & Investment Banking Technology (CCIBT) organization, shaping strategy for infrastructure platforms with enterprise-level influence. As a Business SRE, you will serve as a trusted advisor to senior leadership, architecting automation, observability, and incident management frameworks. You will lead a team of engineers across hybrid cloud environments, driving SRE adoption and governance. This role stands out for its direct impact on Wells Fargo's core banking operations and the chance to define reliability standards enterprise-wide.
Wells Fargo
VerifiedLead Systems Operations Engineer - AWS Certified
Wells Fargo seeks a Lead Systems Operations Engineer to drive reliability and operational excellence across critical platforms. This senior role focuses on high-level systems consultation and large-scale infrastructure planning. Responsibilities include decision-making on technical changes and collaboration with engineering teams to achieve organizational goals.
Wells Fargo
VerifiedPrincipal Engineer, Secure Network Services
This role owns the strategy, architecture, and delivery of secure network solutions across Wells Fargo data centers and Secure Network Interconnect sites. You will lead end-to-end lifecycle ownership from planning through integration of next-generation secure network technologies. As a Principal Engineer, you will influence senior stakeholders and mentor engineers to elevate technical depth and operational maturity. This position drives measurable improvements in network reliability and service resilience at enterprise scale.
Cisco Systems
VerifiedStaff Site Reliability Engineer SRE Cisco Systems Hybrid
Lead technical roadmap for platform reliability scalability and operational excellence. Drive architecture evolution for cloud and air-gapped environments. Establish SLOs and resilience reviews. Lead major initiatives across Kubernetes databases and networking. Mentor engineers and collaborate with cross-functional teams to deliver secure highly reliable deployments.

SRI Tech Solutions
VerifiedLead Site Reliability Engineer SRE
Own the design and scaling of high availability cloud infrastructure for a Generative AI platform. Drive reliability and operational excellence while leading technical initiatives. Shape architecture and mentor engineers across SRE practices.

Intraedge
VerifiedSenior Lead Site Reliability Engineer Principal Engineer Intraedge Phoenix
We seek a senior technical leader to drive reliability and scalability for cloud-native applications. This role focuses on enhancing observability and operational readiness. The ideal candidate will lead initiatives that improve deployment safety and performance.