NinjaOneVerified Source

Senior Site Reliability Engineer, AWS & Observability

160K–240K
Partially · Baltimore, Maryland
Posted August 12, 2026
payroll

Overview

Senior Site Reliability Engineer owns reliability, scalability, and observability for NinjaOne's endpoint management platform, serving 30,000+ customers and millions of endpoints. You join the SRE team in Platform Engineering, collaborating with software engineers who build on Java, Kotlin, C++, Golang, and Postgres running on AWS. You drive improvements in availability, performance, and security through automation and industry-best observability tools. This role stands out for its direct influence on architecture decisions and a 24x7 on-call rotation with real production impact.

What You'll Do7

  • 1Diagnose and resolve complex application and infrastructure issues in AWS environments.
  • 2Participate in the 24x7 on-call rotation, SCRUM, and deployment planning to maintain high service availability.
  • 3Perform Root Cause Analysis (RCA) and provide recommendations for application teams to prevent recurrence.
  • 4Improve availability and reduce customer impact using New Relic, Splunk, and DataDog.
  • 5Influence design decisions to ensure best-practice and security-minded architecture.
  • 6Create and maintain technical documentation and standard operating procedures (SOPs).
  • 7Develop software, scripts, or tooling to improve efficiency and reduce delivery time for applications and infrastructure.

Requirements9

  • 110+ years in DevOps and/or Site Reliability Engineering roles.
  • 2Intermediate+ level Linux administration, scripting, and troubleshooting.
  • 3Demonstrable knowledge of observability tools: New Relic, Splunk, DataDog.
  • 4Comprehensive experience with AWS core services: VPC, EC2, ECS, Route53, Fargate, ALB/NLB.
  • 5Experience with cloud automation and infrastructure-as-code: CloudFormation, Terraform, Helm, Ansible; CDK a plus.
  • 6Good understanding of containers, Fargate, Kubernetes, and distributed microservice architectures.
  • 7Passionate about automation, security, and self-service environments/portals.
  • 8Hands-on experience with CI/CD and SDLC processes.
  • 9Effective communication skills, verbal and written.

Salary Insight

$160 - $240k per year

Location

Typepartially
LocationBaltimore, Maryland

Required Skills

Linux administrationObservability toolsAWSCloudFormationTerraformHelmAnsibleKubernetesCI/CDScriptingAutomation
Share:

Similar open positions

Explore active roles that match your skills and interests.

NinjaOne

11h agoRaleigh, North Carolinapayroll

Senior Site Reliability Engineer, AWS & Observability

You will own the reliability and scalability of NinjaOne's platform, serving 30,000+ customers and millions of endpoints. You will join the SRE team in Platform Engineering, collaborating with engineers who build with Java, Kotlin, Golang, and Postgres on AWS. Your work will directly reduce customer impact and improve service availability. This role stands out with a 24x7 on-call rotation and a chance to influence architecture at scale.

160K–240K
Linux administrationObservability toolsAWS+7 more

NinjaOne

16h agoBaltimore, Marylandpayroll

Staff DevOps Engineer NinjaOne

Staff DevOps Engineer at NinjaOne leading release and deployment pipelines ensuring code moves from commit to production quickly safely and predictably at scale. Own design operation and continuous improvement of NinjaOne's release and deployment pipelines. Collaborate with cross-functional teams to build automation tooling and processes for shipping confidently and often.

150K–230K
AWSCloudFormationCDK+8 more
Xoriant Corporation

Xoriant Corporation

16h agoSan Jose, Californiapayroll

Senior/Staff SRE, AI/ML Platform Infrastructure

Own the reliability of a large-scale AI/ML platform serving millions of requests daily. Kubernetes and Docker are your primary tools. Join a team of 8 SREs supporting 50+ microservices on AWS and GCP. This role focuses on incident command, automation, and platform improvements.

Competitive salary
Production on-callIncident commandBlameless postmortem+15 more
PRIMUS Global Services Inc.

PRIMUS Global Services Inc.

7h agoChicago, Illinoiscontract

Site Reliability Engineer Hybrid

Own and scale high availability production systems across hybrid environments. Lead development of monitoring solutions and automation scripts. Ensure system resilience and performance optimization. Collaborate with cross functional teams to drive reliability initiatives. Differentiate by focusing on real world SRE practices.

45K–50K
pythonawskubernetes+2 more
Intraedge

Intraedge

7h agoPhoenix, Arizonacontract

Senior Lead Site Reliability Engineer Principal Engineer Intraedge Phoenix

We seek a senior technical leader to drive reliability and scalability for cloud-native applications. This role focuses on enhancing observability and operational readiness. The ideal candidate will lead initiatives that improve deployment safety and performance.

Competitive salary
reliabilityobservabilityscalability+3 more
SRI Tech Solutions

SRI Tech Solutions

13h agoOrlando, Floridapayroll

Lead Site Reliability Engineer SRE

Own the design and scaling of high availability cloud infrastructure for a Generative AI platform. Drive reliability and operational excellence while leading technical initiatives. Shape architecture and mentor engineers across SRE practices.

Competitive salary
Site Reliability EngineeringCloud InfrastructureKubernetes+3 more