
Senior/Staff SRE, AI/ML Platform Infrastructure
Overview
Own the reliability of a large-scale AI/ML platform serving millions of requests daily. Kubernetes and Docker are your primary tools. Join a team of 8 SREs supporting 50+ microservices on AWS and GCP. This role focuses on incident command, automation, and platform improvements.
What You'll Do8
- 1Design and implement scalable Kubernetes clusters for AI/ML workloads, ensuring high availability and fault tolerance.
- 2Lead incident response for production outages, following blameless postmortem practices and driving root cause analysis.
- 3Build and maintain infrastructure as code using Terraform or OpenTofu to manage cloud resources on AWS and GCP.
- 4Develop and enhance Prometheus and Grafana dashboards to monitor system health and performance metrics.
- 5Automate deployment and operational tasks using Python or Go to reduce manual intervention and human error.
- 6Collaborate with development teams to improve service reliability through capacity planning and performance tuning.
- 7Drive continuous improvement of the on-call process, including alerting, runbooks, and incident management workflows.
- 8Mentor junior engineers on SRE best practices, containerization, and cloud-native technologies.
Requirements8
- 15+ years of experience in site reliability engineering or similar role, with production on-call experience and incident command.
- 23+ years of hands-on experience with Kubernetes and Docker in production environments.
- 33+ years of experience with cloud platforms like AWS, GCP, or Azure, including infrastructure management.
- 42+ years of experience with Terraform or OpenTofu for infrastructure as code.
- 5Strong knowledge of Prometheus and Grafana for monitoring and alerting.
- 6Experience with CI/CD pipelines and configuration management tools such as Ansible or Puppet.
- 7Solid scripting skills in Python or Go for automation and tooling.
- 8Excellent communication and collaboration skills, with a blameless culture mindset.
Salary Insight
Salary not disclosed in listing
Location
Required Skills
Similar open positions
Explore active roles that match your skills and interests.
Uipath
VerifiedSenior Software Engineer SRE
Design and engineer SRE platform systems using AI while leading cross‑team initiatives. You will identify gaps across teams, design solutions, build, ship, and adopt them, and drive measurable improvements in reliability scalability and performance. This role focuses on ownership accountability and fostering a culture of continuous iteration without relying on generic statements.

SRI Tech Solutions
VerifiedLead Site Reliability Engineer SRE
Own the design and scaling of high availability cloud infrastructure for a Generative AI platform. Drive reliability and operational excellence while leading technical initiatives. Shape architecture and mentor engineers across SRE practices.

PRIMUS Global Services Inc.
VerifiedSite Reliability Engineer Hybrid
Own and scale high availability production systems across hybrid environments. Lead development of monitoring solutions and automation scripts. Ensure system resilience and performance optimization. Collaborate with cross functional teams to drive reliability initiatives. Differentiate by focusing on real world SRE practices.

SilverSearch, Inc.
VerifiedDevOps SRE AI Platform SilverSearch Inc Philadelphia
We seek a skilled DevOps Site Reliability Engineer to build scalable infrastructure for AI workloads at SilverSearch. This role partners with security and engineering teams to operationalize an AI-driven security platform identifying code vulnerabilities and infrastructure risks. The ideal candidate enjoys solving complex challenges and driving innovation.
Hi Marley
VerifiedSr. Software Engineer II DevOps
Lead design and operation of cloud infrastructure on AWS to support core SaaS platform and agentic AI services. Build AI/ML infrastructure and monitoring for LLM-powered services. Establish IaC standards using Terraform. Implement observability beyond availability. Support big data pipelines, warehousing, and analytics. Drive disaster recovery and improve infrastructure parity. Lead architecture reviews and innovate on developer experience. This role differs by focusing on scaling autonomous AI agents in regulated insurance workflows.
Cisco Systems
VerifiedStaff Site Reliability Engineer SRE Cisco Systems Hybrid
Lead technical roadmap for platform reliability scalability and operational excellence. Drive architecture evolution for cloud and air-gapped environments. Establish SLOs and resilience reviews. Lead major initiatives across Kubernetes databases and networking. Mentor engineers and collaborate with cross-functional teams to deliver secure highly reliable deployments.