Senior Site Reliability Engineer, Storage Systems
Overview
You will own the reliability, performance, and capacity of Lambda's production storage fleet across all data centers, operating behind our software-defined data plane. This role sits on the Storage Engineering team, which powers the most demanding AI compute workloads in the industry. You will collaborate with Storage Engineers, Fleet Orchestration, and Release Engineering to automate deployment and configuration. The role demands hands-on troubleshooting of storage incidents, from a single flapping NIC to cluster-wide rebuilds. You will build monitoring, alerting, and self-healing automation to reduce manual toil. What makes this role unique is the scale: tens of thousands of customers and the need for constant innovation in AI infrastructure.
What You'll Do9
- 1Own the reliability, performance, and capacity health of Lambda's production storage fleet, ensuring uptime and efficiency across all data centers.
- 2Build and maintain monitoring, dashboards, and alerting for storage performance, capacity, and hardware failures using Prometheus, Grafana, and Alertmanager.
- 3Investigate and resolve storage-related incidents using deep telemetry, logs, and performance profiling, from a single flapping NIC to a cluster-wide rebuild.
- 4Automate ticketing, escalation, and incident-response workflows to reduce repetitive triage and focus on root cause analysis.
- 5Design and maintain self-healing automation for common failure modes: drive replacement, node swaps, rebuild monitoring, and capacity rebalancing.
- 6Implement CI/CD pipelines for storage automation and tooling using GitHub Actions, Jenkins, and BuildKite.
- 7Partner with Storage Engineers, Fleet Orchestration, and Release Engineering to automate deployment and configuration of software-defined storage.
- 8Work with hardware and networking teams to diagnose low-level I/O and network issues, including NIC errors, RDMA/RoCE/InfiniBand fabric health, and path multipathing.
- 9Participate in an on-call rotation supporting the storage fleet, driving down MTTR and building automation to prevent repeat pages.
Requirements8
- 15+ years operating Linux systems in production or HPC environments, with hands-on storage experience on scale-out or software-defined platforms like CEPH, Lustre, or GPFS.
- 2Hands-on experience operating Software-Defined Storage (SDS) platforms at scale, including integrating with their management and data-plane APIs.
- 3Strong incident-response instincts: comfortable owning a production storage incident end to end, from first alert through root cause to postmortem.
- 4Working experience with monitoring and logging platforms such as Prometheus, Grafana, Alertmanager, Datadog, or SumoLogic, including building dashboards and alert routing.
- 5Working experience with Kubernetes and GitOps tooling such as ArgoCD, Helm/Kustomize, and hands-on troubleshooting.
- 6Working experience with CI/CD tooling (GitHub Actions, Jenkins, BuildKite), containerization (Docker/Podman), and systems programming in Python or Go.
- 7Working experience with Infrastructure as Code (Terraform, Ansible).
- 8Solid understanding of core storage protocols across file (NFS, SMB), object (S3), block (NVMe-oF/TCP), and structured (vector DB, SQL) storage.
Salary Insight
$267 - $356k per year
Location
Required Skills
Similar open positions
Explore active roles that match your skills and interests.

Xoriant Corporation
VerifiedSenior/Staff SRE, AI/ML Platform Infrastructure
Own the reliability of a large-scale AI/ML platform serving millions of requests daily. Kubernetes and Docker are your primary tools. Join a team of 8 SREs supporting 50+ microservices on AWS and GCP. This role focuses on incident command, automation, and platform improvements.
ASK Consulting
VerifiedSenior DevOps Engineer, HPC & EDA Infrastructure
This senior role owns HPC and EDA platform operations, including SLURM cluster administration and datacenter migrations, at an enterprise scale. You will join the IT Datacenter (ITDC) team, coordinating with storage, EDA, and identity teams to deliver critical infrastructure changes. Your work ensures seamless service continuity for semiconductor and scientific computing workloads. This engagement offers direct impact on high-stakes migrations and production stability.

QUANTUM TECHNOLOGIES LLC
VerifiedSenior Linux SRE Infrastructure Engineer
Own infrastructure solutions at QUANTUM TECHNOLOGIES LLC delivering high availability systems in Chandler AZ. Lead scaling and debugging complex environments. Differentiate by driving automation and reliability across cloud native platforms.

Vertisystem Inc.
VerifiedSenior Cloud Storage Solutions Architect
Design and implement high performance enterprise storage solutions across multi vendor platforms and multi cloud providers. Lead architecture and scaling of data backup and recovery for diverse on prem workloads. Drive optimization of storage performance and reliability.

SRI Tech Solutions
VerifiedLead Site Reliability Engineer SRE
Own the design and scaling of high availability cloud infrastructure for a Generative AI platform. Drive reliability and operational excellence while leading technical initiatives. Shape architecture and mentor engineers across SRE practices.

Agama Solutions Inc.
VerifiedSenior Linux SRE Infrastructure Engineer
Lead the administration monitoring and performance tuning of Oracle Enterprise Linux environments at Agama Solutions Inc. Phoenix Arizona onsite contract. Own the design build and lifecycle management of Linux servers storage virtualization and associated infrastructure. Ensure minimal downtime through high availability configurations clustering and load balanced environments. Drive capacity planning performance optimization. This is a unique opportunity to shape platform engineering and operations for a large scale enterprise.