
Senior Site Reliability Engineer, Cloud & Observability
Overview
Own the reliability of distributed systems at scale across multiple locations. You will design, implement, and operate highly available services that power critical customer experiences. You will collaborate with cross functional teams to drive reliability, performance, and efficiency improvements. This role offers the chance to work on cutting edge infrastructure projects in a fast growing company.
What You'll Do5
- 1Build and maintain scalable monitoring and alerting pipelines to reduce mean time to recovery
- 2Design and implement incident response playbooks and post incident reviews to drive continuous improvement
- 3Lead capacity planning, demand forecasting, and cost optimization for multi cloud environments
- 4Ship automated remediation and self healing capabilities to minimize outages
- 5Drive reliability improvements through instrumentation, tracing, and performance tuning
Requirements5
- 15+ years of experience in site reliability engineering, platform engineering, or a related role
- 2Proficiency with cloud environments (e.g., aws, azure, gcp) and container orchestration (e.g., kubernetes)
- 3Strong programming or scripting skills in at least one language such as python, golang, or ruby
- 4Experience building and maintaining observability stacks including metrics, logs, and traces
- 5Strong understanding of networking, security best practices, and incident response
Salary Insight
$100 - $120k per year
Location
Required Skills
Similar open positions
Explore active roles that match your skills and interests.

Intraedge
VerifiedSenior Lead Site Reliability Engineer Principal Engineer Intraedge Phoenix
We seek a senior technical leader to drive reliability and scalability for cloud-native applications. This role focuses on enhancing observability and operational readiness. The ideal candidate will lead initiatives that improve deployment safety and performance.
NinjaOne
VerifiedSenior Site Reliability Engineer, AWS & Observability
Senior Site Reliability Engineer owns reliability, scalability, and observability for NinjaOne's endpoint management platform, serving 30,000+ customers and millions of endpoints. You join the SRE team in Platform Engineering, collaborating with software engineers who build on Java, Kotlin, C++, Golang, and Postgres running on AWS. You drive improvements in availability, performance, and security through automation and industry-best observability tools. This role stands out for its direct influence on architecture decisions and a 24x7 on-call rotation with real production impact.

PRIMUS Global Services Inc.
VerifiedSite Reliability Engineer Hybrid
Own and scale high availability production systems across hybrid environments. Lead development of monitoring solutions and automation scripts. Ensure system resilience and performance optimization. Collaborate with cross functional teams to drive reliability initiatives. Differentiate by focusing on real world SRE practices.
NinjaOne
VerifiedSenior Site Reliability Engineer, AWS & Observability
You will own the reliability and scalability of NinjaOne's platform, serving 30,000+ customers and millions of endpoints. You will join the SRE team in Platform Engineering, collaborating with engineers who build with Java, Kotlin, Golang, and Postgres on AWS. Your work will directly reduce customer impact and improve service availability. This role stands out with a 24x7 on-call rotation and a chance to influence architecture at scale.
Skydio
VerifiedStaff Site Reliability Engineer, Kubernetes & AWS Cloud
Own scalable production infrastructure powering Skydio Cloud platform at scale. Lead reliability for Kubernetes and AWS based systems across multi region deployments. Collaborate with cloud, software, and security teams to ensure availability during critical operations. This role focuses on building and operating a resilient cloud platform with strong emphasis on observability and automation.

Dice Talent & Staffing Solutions
VerifiedSite Reliability Engineer, SRE & Cloud Infrastructure
You will own the reliability and scalability of mission-critical systems for a cutting-edge government contractor in San Francisco. You will design, build, and run large-scale distributed systems with AWS, Kubernetes, and Terraform. You will work with a tight-knit team of engineers to ensure 99.99% uptime for services that matter. This role demands a hands-on engineer who thrives on incident response and automation. You will directly shape the infrastructure strategy from day one.