LambdaVerified Source
Remote

Senior Site Reliability Engineer, Storage Systems

267K–356K
Remote · San Jose, California
Posted August 12, 2026
payroll

Overview

You will own the reliability, performance, and capacity of Lambda's production storage fleet across all data centers, operating behind our software-defined data plane. This role sits on the Storage Engineering team, which powers the most demanding AI compute workloads in the industry. You will collaborate with Storage Engineers, Fleet Orchestration, and Release Engineering to automate deployment and configuration. The role demands hands-on troubleshooting of storage incidents, from a single flapping NIC to cluster-wide rebuilds. You will build monitoring, alerting, and self-healing automation to reduce manual toil. What makes this role unique is the scale: tens of thousands of customers and the need for constant innovation in AI infrastructure.

What You'll Do9

  • 1Own the reliability, performance, and capacity health of Lambda's production storage fleet, ensuring uptime and efficiency across all data centers.
  • 2Build and maintain monitoring, dashboards, and alerting for storage performance, capacity, and hardware failures using Prometheus, Grafana, and Alertmanager.
  • 3Investigate and resolve storage-related incidents using deep telemetry, logs, and performance profiling, from a single flapping NIC to a cluster-wide rebuild.
  • 4Automate ticketing, escalation, and incident-response workflows to reduce repetitive triage and focus on root cause analysis.
  • 5Design and maintain self-healing automation for common failure modes: drive replacement, node swaps, rebuild monitoring, and capacity rebalancing.
  • 6Implement CI/CD pipelines for storage automation and tooling using GitHub Actions, Jenkins, and BuildKite.
  • 7Partner with Storage Engineers, Fleet Orchestration, and Release Engineering to automate deployment and configuration of software-defined storage.
  • 8Work with hardware and networking teams to diagnose low-level I/O and network issues, including NIC errors, RDMA/RoCE/InfiniBand fabric health, and path multipathing.
  • 9Participate in an on-call rotation supporting the storage fleet, driving down MTTR and building automation to prevent repeat pages.

Requirements8

  • 15+ years operating Linux systems in production or HPC environments, with hands-on storage experience on scale-out or software-defined platforms like CEPH, Lustre, or GPFS.
  • 2Hands-on experience operating Software-Defined Storage (SDS) platforms at scale, including integrating with their management and data-plane APIs.
  • 3Strong incident-response instincts: comfortable owning a production storage incident end to end, from first alert through root cause to postmortem.
  • 4Working experience with monitoring and logging platforms such as Prometheus, Grafana, Alertmanager, Datadog, or SumoLogic, including building dashboards and alert routing.
  • 5Working experience with Kubernetes and GitOps tooling such as ArgoCD, Helm/Kustomize, and hands-on troubleshooting.
  • 6Working experience with CI/CD tooling (GitHub Actions, Jenkins, BuildKite), containerization (Docker/Podman), and systems programming in Python or Go.
  • 7Working experience with Infrastructure as Code (Terraform, Ansible).
  • 8Solid understanding of core storage protocols across file (NFS, SMB), object (S3), block (NVMe-oF/TCP), and structured (vector DB, SQL) storage.

Salary Insight

$267 - $356k per year

Location

Typeremote
LocationSan Jose, California
This is a remote position

Required Skills

pythongokubernetesterraformansible
Share:

Similar open positions

Explore active roles that match your skills and interests.

Xoriant Corporation

Xoriant Corporation

22h agoSan Jose, Californiapayroll

Senior/Staff SRE, AI/ML Platform Infrastructure

Own the reliability of a large-scale AI/ML platform serving millions of requests daily. Kubernetes and Docker are your primary tools. Join a team of 8 SREs supporting 50+ microservices on AWS and GCP. This role focuses on incident command, automation, and platform improvements.

Competitive salary
Production on-callIncident commandBlameless postmortem+15 more

ASK Consulting

16h agoSacramento, Californiapayroll

Senior DevOps Engineer, HPC & EDA Infrastructure

This senior role owns HPC and EDA platform operations, including SLURM cluster administration and datacenter migrations, at an enterprise scale. You will join the IT Datacenter (ITDC) team, coordinating with storage, EDA, and identity teams to deliver critical infrastructure changes. Your work ensures seamless service continuity for semiconductor and scientific computing workloads. This engagement offers direct impact on high-stakes migrations and production stability.

125K–133K
slurmslesansible+2 more
QUANTUM TECHNOLOGIES LLC

QUANTUM TECHNOLOGIES LLC

19h agoPhoenix, Arizonacontract

Senior Linux SRE Infrastructure Engineer

Own infrastructure solutions at QUANTUM TECHNOLOGIES LLC delivering high availability systems in Chandler AZ. Lead scaling and debugging complex environments. Differentiate by driving automation and reliability across cloud native platforms.

62K
Oracle Enterprise LinuxHigh AvailabilityClustering+5 more
Vertisystem Inc.

Vertisystem Inc.

21h agoLos Angeles, Californiacontract

Senior Cloud Storage Solutions Architect

Design and implement high performance enterprise storage solutions across multi vendor platforms and multi cloud providers. Lead architecture and scaling of data backup and recovery for diverse on prem workloads. Drive optimization of storage performance and reliability.

70K–87K
DesigningImplementingOptimizing high-performance enterprise storage solutions+8 more
SRI Tech Solutions

SRI Tech Solutions

19h agoOrlando, Floridapayroll

Lead Site Reliability Engineer SRE

Own the design and scaling of high availability cloud infrastructure for a Generative AI platform. Drive reliability and operational excellence while leading technical initiatives. Shape architecture and mentor engineers across SRE practices.

Competitive salary
Site Reliability EngineeringCloud InfrastructureKubernetes+3 more
Agama Solutions Inc.

Agama Solutions Inc.

19h agoPhoenix, Arizonacontract

Senior Linux SRE Infrastructure Engineer

Lead the administration monitoring and performance tuning of Oracle Enterprise Linux environments at Agama Solutions Inc. Phoenix Arizona onsite contract. Own the design build and lifecycle management of Linux servers storage virtualization and associated infrastructure. Ensure minimal downtime through high availability configurations clustering and load balanced environments. Drive capacity planning performance optimization. This is a unique opportunity to shape platform engineering and operations for a large scale enterprise.

60K–65K
Oracle Enterprise LinuxHigh AvailabilityClustering+5 more