Senior Site Reliability Engineer, AWS & Observability
Overview
You will own the reliability and scalability of NinjaOne's platform, serving 30,000+ customers and millions of endpoints. You will join the SRE team in Platform Engineering, collaborating with engineers who build with Java, Kotlin, Golang, and Postgres on AWS. Your work will directly reduce customer impact and improve service availability. This role stands out with a 24x7 on-call rotation and a chance to influence architecture at scale.
What You'll Do9
- 1Diagnose and resolve complex application and infrastructure issues using New Relic, Splunk, and DataDog.
- 2Participate in the 24x7 on-call rotation and deployment planning to ensure high availability.
- 3Perform root cause analysis (RCA) and provide actionable recommendations to application teams.
- 4Improve availability and reduce customer impact by implementing industry-best observability practices.
- 5Influence design decisions to ensure best-practice and security-minded architecture.
- 6Create and maintain technical documentation and standard operating procedures (SOPs).
- 7Develop software, scripts, or tooling to automate and improve delivery efficiency.
- 8Collaborate with cross-functional teams to scale AWS infrastructure using CloudFormation, Terraform, and Ansible.
- 9Ensure security and compliance within the FedRAMP-authorized environment.
Requirements10
- 110+ years in DevOps or Site Reliability Engineering roles.
- 2Intermediate+ Linux administration, scripting, and troubleshooting skills.
- 3Hands-on experience with observability tools: New Relic, Splunk, DataDog.
- 4Comprehensive experience with AWS core services: VPC, EC2, ECS, Route53, Fargate, ALB/NLB.
- 5Experience with infrastructure-as-code: CloudFormation, Terraform, Helm, Ansible; CDK is a plus.
- 6Good understanding of containers, Fargate, Kubernetes, and distributed microservice architecture.
- 7Passion for automation, security, and self-service environments.
- 8Hands-on experience with CI/CD and SDLC processes.
- 9Effective verbal and written communication skills.
- 10Must be a U.S. citizen or lawful permanent resident due to federal security requirements.
Salary Insight
$160 - $240k per year
Location
Required Skills
Similar open positions
Explore active roles that match your skills and interests.
NinjaOne
VerifiedSenior Site Reliability Engineer, AWS & Observability
Senior Site Reliability Engineer owns reliability, scalability, and observability for NinjaOne's endpoint management platform, serving 30,000+ customers and millions of endpoints. You join the SRE team in Platform Engineering, collaborating with software engineers who build on Java, Kotlin, C++, Golang, and Postgres running on AWS. You drive improvements in availability, performance, and security through automation and industry-best observability tools. This role stands out for its direct influence on architecture decisions and a 24x7 on-call rotation with real production impact.

Xoriant Corporation
VerifiedSenior/Staff SRE, AI/ML Platform Infrastructure
Own the reliability of a large-scale AI/ML platform serving millions of requests daily. Kubernetes and Docker are your primary tools. Join a team of 8 SREs supporting 50+ microservices on AWS and GCP. This role focuses on incident command, automation, and platform improvements.
NinjaOne
VerifiedStaff DevOps Engineer NinjaOne
Staff DevOps Engineer at NinjaOne leading release and deployment pipelines ensuring code moves from commit to production quickly safely and predictably at scale. Own design operation and continuous improvement of NinjaOne's release and deployment pipelines. Collaborate with cross-functional teams to build automation tooling and processes for shipping confidently and often.

PRIMUS Global Services Inc.
VerifiedSite Reliability Engineer Hybrid
Own and scale high availability production systems across hybrid environments. Lead development of monitoring solutions and automation scripts. Ensure system resilience and performance optimization. Collaborate with cross functional teams to drive reliability initiatives. Differentiate by focusing on real world SRE practices.

Intraedge
VerifiedSenior Lead Site Reliability Engineer Principal Engineer Intraedge Phoenix
We seek a senior technical leader to drive reliability and scalability for cloud-native applications. This role focuses on enhancing observability and operational readiness. The ideal candidate will lead initiatives that improve deployment safety and performance.
C0035 LiveRamp, Inc.
VerifiedSenior SRE (Site Reliability Engineer) - LiveRamp
You will own the deployment and reliability of global products at LiveRamp, the data collaboration platform used by hundreds of innovative companies. You will set up production and internal environments, provide 24/7 first-line support, and drive resolutions with engineering teams. You will enhance CI/CD tooling, maintain Terraform scripts, and optimize system performance and cost. This role stands out for its global scope, working with teams across California, Paris, Nantong, Singapore, and Australia, and for championing SRE best practices.