
Site Reliability Engineer Hybrid
Overview
Own and scale high availability production systems across hybrid environments. Lead development of monitoring solutions and automation scripts. Ensure system resilience and performance optimization. Collaborate with cross functional teams to drive reliability initiatives. Differentiate by focusing on real world SRE practices.
What You'll Do10
- 1Build and maintain observability dashboards using Prometheus Grafana
- 2Design and implement automated deployment pipelines with Jenkins
- 3Develop machine learning models for predictive maintenance
- 4Scale Kubernetes clusters using AWS EKS
- 5Debug complex distributed system failures
- 6Drive incident response and postmortem analysis
- 7Optimize database query performance with indexing strategies
- 8Implement CI/CD workflows for microservices architecture
- 9Monitor system metrics and alert thresholds
- 10Collaborate with product owners to define SLIs and SLOs
Requirements10
- 15+ years building SRE solutions with Python and AWS
- 23-5 seasons managing production environments
- 3Expertise in Kubernetes and Docker orchestration
- 4Proficiency in React for internal dashboards
- 5Strong understanding of machine learning libraries like scikit-learn
- 6Experience with Prometheus Grafana stack
- 7Familiarity with Jenkins continuous integration
- 8Knowledge of cloud native technologies such as AWS
- 9Ability to write scalable code in Java or Go
- 10Certification in AWS Certified Solutions Architect
Salary Insight
$45 - $50k per year
Location
Required Skills
Similar open positions
Explore active roles that match your skills and interests.

SRI Tech Solutions
VerifiedLead Site Reliability Engineer SRE
Own the design and scaling of high availability cloud infrastructure for a Generative AI platform. Drive reliability and operational excellence while leading technical initiatives. Shape architecture and mentor engineers across SRE practices.

Info Way Solutions
VerifiedSRE Architect Info Way Solutions Seattle
Lead enterprise-wide reliability engineering, observability, and operational excellence initiatives. Drive SRE transformation define reliability strategies establish SLO governance and lead reliability engineering adoption across large-scale enterprise environments.

Xoriant Corporation
VerifiedSenior/Staff SRE, AI/ML Platform Infrastructure
Own the reliability of a large-scale AI/ML platform serving millions of requests daily. Kubernetes and Docker are your primary tools. Join a team of 8 SREs supporting 50+ microservices on AWS and GCP. This role focuses on incident command, automation, and platform improvements.
Uipath
VerifiedSenior Software Engineer SRE
Design and engineer SRE platform systems using AI while leading cross‑team initiatives. You will identify gaps across teams, design solutions, build, ship, and adopt them, and drive measurable improvements in reliability scalability and performance. This role focuses on ownership accountability and fostering a culture of continuous iteration without relying on generic statements.

BCforward
VerifiedTechnology & Information Architectures - Performance & Reliability Engineer
Own performance and reliability engineering initiatives scaling across hybrid cloud environments. Deliver high availability solutions leveraging Python AWS React Kubernetes and Spark. Drive observability and automation using industry best practices. Impact multiple global teams.
Cisco Systems
VerifiedStaff Site Reliability Engineer SRE Cisco Systems Hybrid
Lead technical roadmap for platform reliability scalability and operational excellence. Drive architecture evolution for cloud and air-gapped environments. Establish SLOs and resilience reviews. Lead major initiatives across Kubernetes databases and networking. Mentor engineers and collaborate with cross-functional teams to deliver secure highly reliable deployments.