
SRE Lead, AWS & Kubernetes (Onsite)
Overview
You own reliability for a global eCommerce platform, defining SLIs, SLOs, and error budgets across AWS and Kubernetes. You lead a team of SREs and drive chaos engineering to validate fault tolerance at scale. You architect multi-AZ, multi-region deployments for zero-downtime operations, and you own the CI/CD pipelines on Jenkins and EKS. This role combines deep technical leadership with hands-on engineering, impacting millions of users daily.
What You'll Do9
- 1Define, monitor, and enforce SLIs, SLOs, and error budgets for all eCommerce services.
- 2Drive chaos engineering and game-day exercises to validate fault tolerance.
- 3Architect resilient multi-AZ, multi-region AWS deployments for zero-downtime operations.
- 4Own and evolve Jenkins-based CI/CD pipelines for microservices on Kubernetes/EKS.
- 5Implement GitOps workflows and enforce trunk-based development.
- 6Lead incident response and postmortems, driving systemic improvements.
- 7Collaborate with development teams to improve service reliability and performance.
- 8Automate infrastructure provisioning and configuration management with Terraform and Ansible.
- 9Design and implement monitoring and alerting using Prometheus, Grafana, and Datadog.
Requirements9
- 15+ years of SRE or DevOps experience, with 2+ years in a leadership role.
- 25+ years building and scaling on AWS, including EC2, S3, VPC, and IAM.
- 33+ years managing Kubernetes clusters in production, with EKS experience.
- 45+ years with Jenkins or equivalent CI/CD tools, and Git workflows including GitOps.
- 53+ years coding in Python, Go, or Java for automation and tooling.
- 6Deep understanding of SLIs, SLOs, and error budgets, with practical implementation experience.
- 7Proven track record of driving chaos engineering and game-day exercises.
- 8Strong background in infrastructure as code using Terraform and Ansible.
- 9Excellent communication and collaboration skills, with experience leading on-call rotations
Salary Insight
Salary not disclosed in listing
Location
Required Skills
Similar open positions
Explore active roles that match your skills and interests.

Xoriant Corporation
VerifiedSenior/Staff SRE, AI/ML Platform Infrastructure
Own the reliability of a large-scale AI/ML platform serving millions of requests daily. Kubernetes and Docker are your primary tools. Join a team of 8 SREs supporting 50+ microservices on AWS and GCP. This role focuses on incident command, automation, and platform improvements.

SRI Tech Solutions
VerifiedLead Site Reliability Engineer SRE
Own the design and scaling of high availability cloud infrastructure for a Generative AI platform. Drive reliability and operational excellence while leading technical initiatives. Shape architecture and mentor engineers across SRE practices.
Uipath
VerifiedSenior Software Engineer SRE
Design and engineer SRE platform systems using AI while leading cross‑team initiatives. You will identify gaps across teams, design solutions, build, ship, and adopt them, and drive measurable improvements in reliability scalability and performance. This role focuses on ownership accountability and fostering a culture of continuous iteration without relying on generic statements.

PRIMUS Global Services Inc.
VerifiedSite Reliability Engineer Hybrid
Own and scale high availability production systems across hybrid environments. Lead development of monitoring solutions and automation scripts. Ensure system resilience and performance optimization. Collaborate with cross functional teams to drive reliability initiatives. Differentiate by focusing on real world SRE practices.

Info Way Solutions
VerifiedSRE Architect Info Way Solutions Seattle
Lead enterprise-wide reliability engineering, observability, and operational excellence initiatives. Drive SRE transformation define reliability strategies establish SLO governance and lead reliability engineering adoption across large-scale enterprise environments.

Dice Talent & Staffing Solutions
VerifiedSite Reliability Engineer, SRE & Cloud Infrastructure
You will own the reliability and scalability of mission-critical systems for a cutting-edge government contractor in San Francisco. You will design, build, and run large-scale distributed systems with AWS, Kubernetes, and Terraform. You will work with a tight-knit team of engineers to ensure 99.99% uptime for services that matter. This role demands a hands-on engineer who thrives on incident response and automation. You will directly shape the infrastructure strategy from day one.