Senior Site Reliability Engineer, AWS & Observability
Overview
Senior Site Reliability Engineer owns reliability, scalability, and observability for NinjaOne's endpoint management platform, serving 30,000+ customers and millions of endpoints. You join the SRE team in Platform Engineering, collaborating with software engineers who build on Java, Kotlin, C++, Golang, and Postgres running on AWS. You drive improvements in availability, performance, and security through automation and industry-best observability tools. This role stands out for its direct influence on architecture decisions and a 24x7 on-call rotation with real production impact.
What You'll Do7
- 1Diagnose and resolve complex application and infrastructure issues in AWS environments.
- 2Participate in the 24x7 on-call rotation, SCRUM, and deployment planning to maintain high service availability.
- 3Perform Root Cause Analysis (RCA) and provide recommendations for application teams to prevent recurrence.
- 4Improve availability and reduce customer impact using New Relic, Splunk, and DataDog.
- 5Influence design decisions to ensure best-practice and security-minded architecture.
- 6Create and maintain technical documentation and standard operating procedures (SOPs).
- 7Develop software, scripts, or tooling to improve efficiency and reduce delivery time for applications and infrastructure.
Requirements9
- 110+ years in DevOps and/or Site Reliability Engineering roles.
- 2Intermediate+ level Linux administration, scripting, and troubleshooting.
- 3Demonstrable knowledge of observability tools: New Relic, Splunk, DataDog.
- 4Comprehensive experience with AWS core services: VPC, EC2, ECS, Route53, Fargate, ALB/NLB.
- 5Experience with cloud automation and infrastructure-as-code: CloudFormation, Terraform, Helm, Ansible; CDK a plus.
- 6Good understanding of containers, Fargate, Kubernetes, and distributed microservice architectures.
- 7Passionate about automation, security, and self-service environments/portals.
- 8Hands-on experience with CI/CD and SDLC processes.
- 9Effective communication skills, verbal and written.
Salary Insight
$160 - $240k per year
Location
Required Skills
Similar open positions
Explore active roles that match your skills and interests.
NinjaOne
VerifiedSenior Site Reliability Engineer, AWS & Observability
You will own the reliability and scalability of NinjaOne's platform, serving 30,000+ customers and millions of endpoints. You will join the SRE team in Platform Engineering, collaborating with engineers who build with Java, Kotlin, Golang, and Postgres on AWS. Your work will directly reduce customer impact and improve service availability. This role stands out with a 24x7 on-call rotation and a chance to influence architecture at scale.
NinjaOne
VerifiedStaff DevOps Engineer NinjaOne
Staff DevOps Engineer at NinjaOne leading release and deployment pipelines ensuring code moves from commit to production quickly safely and predictably at scale. Own design operation and continuous improvement of NinjaOne's release and deployment pipelines. Collaborate with cross-functional teams to build automation tooling and processes for shipping confidently and often.

Xoriant Corporation
VerifiedSenior/Staff SRE, AI/ML Platform Infrastructure
Own the reliability of a large-scale AI/ML platform serving millions of requests daily. Kubernetes and Docker are your primary tools. Join a team of 8 SREs supporting 50+ microservices on AWS and GCP. This role focuses on incident command, automation, and platform improvements.

PRIMUS Global Services Inc.
VerifiedSite Reliability Engineer Hybrid
Own and scale high availability production systems across hybrid environments. Lead development of monitoring solutions and automation scripts. Ensure system resilience and performance optimization. Collaborate with cross functional teams to drive reliability initiatives. Differentiate by focusing on real world SRE practices.

Intraedge
VerifiedSenior Lead Site Reliability Engineer Principal Engineer Intraedge Phoenix
We seek a senior technical leader to drive reliability and scalability for cloud-native applications. This role focuses on enhancing observability and operational readiness. The ideal candidate will lead initiatives that improve deployment safety and performance.

SRI Tech Solutions
VerifiedLead Site Reliability Engineer SRE
Own the design and scaling of high availability cloud infrastructure for a Generative AI platform. Drive reliability and operational excellence while leading technical initiatives. Shape architecture and mentor engineers across SRE practices.