Xoriant Corporation
Xoriant CorporationVerified Source

Senior/Staff SRE, AI/ML Platform Infrastructure

Onsite · San Jose, California
Posted August 12, 2026
payroll

Overview

Own the reliability of a large-scale AI/ML platform serving millions of requests daily. Kubernetes and Docker are your primary tools. Join a team of 8 SREs supporting 50+ microservices on AWS and GCP. This role focuses on incident command, automation, and platform improvements.

What You'll Do8

  • 1Design and implement scalable Kubernetes clusters for AI/ML workloads, ensuring high availability and fault tolerance.
  • 2Lead incident response for production outages, following blameless postmortem practices and driving root cause analysis.
  • 3Build and maintain infrastructure as code using Terraform or OpenTofu to manage cloud resources on AWS and GCP.
  • 4Develop and enhance Prometheus and Grafana dashboards to monitor system health and performance metrics.
  • 5Automate deployment and operational tasks using Python or Go to reduce manual intervention and human error.
  • 6Collaborate with development teams to improve service reliability through capacity planning and performance tuning.
  • 7Drive continuous improvement of the on-call process, including alerting, runbooks, and incident management workflows.
  • 8Mentor junior engineers on SRE best practices, containerization, and cloud-native technologies.

Requirements8

  • 15+ years of experience in site reliability engineering or similar role, with production on-call experience and incident command.
  • 23+ years of hands-on experience with Kubernetes and Docker in production environments.
  • 33+ years of experience with cloud platforms like AWS, GCP, or Azure, including infrastructure management.
  • 42+ years of experience with Terraform or OpenTofu for infrastructure as code.
  • 5Strong knowledge of Prometheus and Grafana for monitoring and alerting.
  • 6Experience with CI/CD pipelines and configuration management tools such as Ansible or Puppet.
  • 7Solid scripting skills in Python or Go for automation and tooling.
  • 8Excellent communication and collaboration skills, with a blameless culture mindset.

Salary Insight

Salary not disclosed in listing

Location

Typeonsite
LocationSan Jose, California

Required Skills

Production on-callIncident commandBlameless postmortemKubernetesDockerCloud-native infrastructureAWSGoogle Cloud PlatformAzureTerraformOpenTofuPrometheusGrafanaMetricsLoggingAlertingDashboard designAlert design
Share:

Similar open positions

Explore active roles that match your skills and interests.

Uipath

19h agoDenver, Coloradopayroll

Senior Software Engineer SRE

Design and engineer SRE platform systems using AI while leading cross‑team initiatives. You will identify gaps across teams, design solutions, build, ship, and adopt them, and drive measurable improvements in reliability scalability and performance. This role focuses on ownership accountability and fostering a culture of continuous iteration without relying on generic statements.

160K–210K
pythonawsreact+2 more
SRI Tech Solutions

SRI Tech Solutions

17h agoOrlando, Floridapayroll

Lead Site Reliability Engineer SRE

Own the design and scaling of high availability cloud infrastructure for a Generative AI platform. Drive reliability and operational excellence while leading technical initiatives. Shape architecture and mentor engineers across SRE practices.

Competitive salary
Site Reliability EngineeringCloud InfrastructureKubernetes+3 more
PRIMUS Global Services Inc.

PRIMUS Global Services Inc.

11h agoChicago, Illinoiscontract

Site Reliability Engineer Hybrid

Own and scale high availability production systems across hybrid environments. Lead development of monitoring solutions and automation scripts. Ensure system resilience and performance optimization. Collaborate with cross functional teams to drive reliability initiatives. Differentiate by focusing on real world SRE practices.

45K–50K
pythonawskubernetes+2 more
SilverSearch, Inc.

SilverSearch, Inc.

15d agoPhiladelphia, Pennsylvaniapayroll

DevOps SRE AI Platform SilverSearch Inc Philadelphia

We seek a skilled DevOps Site Reliability Engineer to build scalable infrastructure for AI workloads at SilverSearch. This role partners with security and engineering teams to operationalize an AI-driven security platform identifying code vulnerabilities and infrastructure risks. The ideal candidate enjoys solving complex challenges and driving innovation.

Competitive salary
DevOpsSite Reliability EngineeringAI platforms+4 more

Hi Marley

19h agoBoston, Massachusettspayroll

Sr. Software Engineer II DevOps

Lead design and operation of cloud infrastructure on AWS to support core SaaS platform and agentic AI services. Build AI/ML infrastructure and monitoring for LLM-powered services. Establish IaC standards using Terraform. Implement observability beyond availability. Support big data pipelines, warehousing, and analytics. Drive disaster recovery and improve infrastructure parity. Lead architecture reviews and innovate on developer experience. This role differs by focusing on scaling autonomous AI agents in regulated insurance workflows.

119K–221K
AWSTerraformPython+4 more
Cisco Systems

Cisco Systems

11d agoSan Francisco, Californiapayroll

Staff Site Reliability Engineer SRE Cisco Systems Hybrid

Lead technical roadmap for platform reliability scalability and operational excellence. Drive architecture evolution for cloud and air-gapped environments. Establish SLOs and resilience reviews. Lead major initiatives across Kubernetes databases and networking. Mentor engineers and collaborate with cross-functional teams to deliver secure highly reliable deployments.

187K–268K
KubernetesAWSGCP+8 more