VLink Inc
VLink IncVerified Source

Senior Site Reliability Engineer, Cloud & Observability

100K–120K
Onsite · Dallas, Texas
Posted August 14, 2026
full-time

Overview

Own the reliability of distributed systems at scale across multiple locations. You will design, implement, and operate highly available services that power critical customer experiences. You will collaborate with cross functional teams to drive reliability, performance, and efficiency improvements. This role offers the chance to work on cutting edge infrastructure projects in a fast growing company.

What You'll Do5

  • 1Build and maintain scalable monitoring and alerting pipelines to reduce mean time to recovery
  • 2Design and implement incident response playbooks and post incident reviews to drive continuous improvement
  • 3Lead capacity planning, demand forecasting, and cost optimization for multi cloud environments
  • 4Ship automated remediation and self healing capabilities to minimize outages
  • 5Drive reliability improvements through instrumentation, tracing, and performance tuning

Requirements5

  • 15+ years of experience in site reliability engineering, platform engineering, or a related role
  • 2Proficiency with cloud environments (e.g., aws, azure, gcp) and container orchestration (e.g., kubernetes)
  • 3Strong programming or scripting skills in at least one language such as python, golang, or ruby
  • 4Experience building and maintaining observability stacks including metrics, logs, and traces
  • 5Strong understanding of networking, security best practices, and incident response

Salary Insight

$100 - $120k per year

Location

Typeonsite
LocationDallas, Texas

Required Skills

pythonawskubernetesobservabilitymonitoring
Share:

Similar open positions

Explore active roles that match your skills and interests.

Intraedge

Intraedge

2d agoPhoenix, Arizonacontract

Senior Lead Site Reliability Engineer Principal Engineer Intraedge Phoenix

We seek a senior technical leader to drive reliability and scalability for cloud-native applications. This role focuses on enhancing observability and operational readiness. The ideal candidate will lead initiatives that improve deployment safety and performance.

Competitive salary
reliabilityobservabilityscalability+3 more

NinjaOne

2d agoBaltimore, Marylandfull-time

Senior Site Reliability Engineer, AWS & Observability

Senior Site Reliability Engineer owns reliability, scalability, and observability for NinjaOne's endpoint management platform, serving 30,000+ customers and millions of endpoints. You join the SRE team in Platform Engineering, collaborating with software engineers who build on Java, Kotlin, C++, Golang, and Postgres running on AWS. You drive improvements in availability, performance, and security through automation and industry-best observability tools. This role stands out for its direct influence on architecture decisions and a 24x7 on-call rotation with real production impact.

160K–240K
Linux administrationObservability toolsAWS+8 more
PRIMUS Global Services Inc.

PRIMUS Global Services Inc.

2d agoChicago, Illinoiscontract

Site Reliability Engineer Hybrid

Own and scale high availability production systems across hybrid environments. Lead development of monitoring solutions and automation scripts. Ensure system resilience and performance optimization. Collaborate with cross functional teams to drive reliability initiatives. Differentiate by focusing on real world SRE practices.

45K–50K
pythonawskubernetes+2 more

NinjaOne

2d agoRaleigh, North Carolinafull-time

Senior Site Reliability Engineer, AWS & Observability

You will own the reliability and scalability of NinjaOne's platform, serving 30,000+ customers and millions of endpoints. You will join the SRE team in Platform Engineering, collaborating with engineers who build with Java, Kotlin, Golang, and Postgres on AWS. Your work will directly reduce customer impact and improve service availability. This role stands out with a 24x7 on-call rotation and a chance to influence architecture at scale.

160K–240K
Linux administrationObservability toolsAWS+7 more

Skydio

9h agoRemotefull-time

Staff Site Reliability Engineer, Kubernetes & AWS Cloud

Own scalable production infrastructure powering Skydio Cloud platform at scale. Lead reliability for Kubernetes and AWS based systems across multi region deployments. Collaborate with cloud, software, and security teams to ensure availability during critical operations. This role focuses on building and operating a resilient cloud platform with strong emphasis on observability and automation.

Competitive salary
kubernetesawsterraform+2 more
Dice Talent & Staffing Solutions

Dice Talent & Staffing Solutions

1d agoSan Francisco, Californiafull-time

Site Reliability Engineer, SRE & Cloud Infrastructure

You will own the reliability and scalability of mission-critical systems for a cutting-edge government contractor in San Francisco. You will design, build, and run large-scale distributed systems with AWS, Kubernetes, and Terraform. You will work with a tight-knit team of engineers to ensure 99.99% uptime for services that matter. This role demands a hands-on engineer who thrives on incident response and automation. You will directly shape the infrastructure strategy from day one.

180K–265K
awskubernetesterraform+2 more