SkydioVerified Source
Remote

Staff Site Reliability Engineer, Kubernetes & AWS Cloud

Remote · San Francisco, California
Posted August 14, 2026
full-time

Overview

Own scalable production infrastructure powering Skydio Cloud platform at scale. Lead reliability for Kubernetes and AWS based systems across multi region deployments. Collaborate with cloud, software, and security teams to ensure availability during critical operations. This role focuses on building and operating a resilient cloud platform with strong emphasis on observability and automation.

What You'll Do10

  • 1Build, operate, and troubleshoot production Kubernetes/EKS clusters and upgrades
  • 2Manage AWS infrastructure including VPCs, subnets, load balancers, IAM, databases
  • 3Define and maintain infrastructure using Terraform and CI/CD pipelines
  • 4Build monitoring, alerting, and observability for critical infrastructure
  • 5Automate operational tasks using Python or Go in production environments
  • 6Participate in on-call rotations and respond to incidents to minimize downtime
  • 7Scale infrastructure across new regions and deployment environments
  • 8Diagnose production issues across Kubernetes, AWS, Linux, networking, and databases
  • 9Own deployment systems such as Argo CD, Spinnaker, GitHub Actions, GitLab CI/CD, or Jenkins
  • 10Improve reliability and performance through proactive problem detection and resolution

Requirements9

  • 18+ years of experience in Site Reliability Engineering or related roles
  • 2Strong hands-on experience operating Kubernetes and production clusters
  • 3Proficiency with AWS fundamentals (VPCs, subnets, networking, load balancers, IAM, EKS)
  • 4Production experience with Terraform or similar IaC tooling
  • 5Experience owning CI/CD and deployment systems (Argo CD, Spinnaker, GitHub Actions, GitLab CI/CD, Jenkins)
  • 6Experience diagnosing production infrastructure and networking problems
  • 7Experience solving meaningful scaling or reliability challenges
  • 8Proficiency automating operational tasks with Python or Go
  • 9Experience with multi-region infrastructure or on-premises deployments (bonus)

Salary Insight

Salary not disclosed in listing

Location

Typeremote
LocationSan Francisco, California
This is a remote position

Required Skills

kubernetesawsterraformci/cdpython
Share:

Similar open positions

Explore active roles that match your skills and interests.

VLink Inc

VLink Inc

10h agoDallas, Texasfull-time

Senior Site Reliability Engineer, Cloud & Observability

Own the reliability of distributed systems at scale across multiple locations. You will design, implement, and operate highly available services that power critical customer experiences. You will collaborate with cross functional teams to drive reliability, performance, and efficiency improvements. This role offers the chance to work on cutting edge infrastructure projects in a fast growing company.

100K–120K
pythonawskubernetes+2 more
System One

System One

1d agoPittsburgh, Pennsylvaniafull-time

Site Reliability Engineer, Production Operations

You own production reliability for critical internal and external applications across Pittsburgh, Cleveland, and Dallas sites. You drive incident response, performance tuning, and automation to maintain high availability. You collaborate with production support and engineering teams to reduce downtime and improve system resilience. This role focuses on proactive monitoring and reliability engineering, not just firefighting.

Competitive salary
pythonawskubernetes+2 more
Dice Talent & Staffing Solutions

Dice Talent & Staffing Solutions

1d agoSan Francisco, Californiafull-time

Site Reliability Engineer, SRE & Cloud Infrastructure

You will own the reliability and scalability of mission-critical systems for a cutting-edge government contractor in San Francisco. You will design, build, and run large-scale distributed systems with AWS, Kubernetes, and Terraform. You will work with a tight-knit team of engineers to ensure 99.99% uptime for services that matter. This role demands a hands-on engineer who thrives on incident response and automation. You will directly shape the infrastructure strategy from day one.

180K–265K
awskubernetesterraform+2 more
SRI Tech Solutions

SRI Tech Solutions

2d agoOrlando, Floridafull-time

Lead Site Reliability Engineer SRE

Own the design and scaling of high availability cloud infrastructure for a Generative AI platform. Drive reliability and operational excellence while leading technical initiatives. Shape architecture and mentor engineers across SRE practices.

Competitive salary
Site Reliability EngineeringCloud InfrastructureKubernetes+3 more
PRIMUS Global Services Inc.

PRIMUS Global Services Inc.

2d agoChicago, Illinoiscontract

Site Reliability Engineer Hybrid

Own and scale high availability production systems across hybrid environments. Lead development of monitoring solutions and automation scripts. Ensure system resilience and performance optimization. Collaborate with cross functional teams to drive reliability initiatives. Differentiate by focusing on real world SRE practices.

45K–50K
pythonawskubernetes+2 more

superblocks

1d agoNew York, New Yorkfull-time

Site Reliability Engineer, Kubernetes & AI Infrastructure

You will own the AWS-backed multi-tenant cloud and on-premise infrastructure that runs hundreds of thousands of AI applications at Superblocks. You will join a team of engineers from Uber, Stripe, and Datadog who built systems like Kafka, Kibana, and Debezium. We are a fast-growing AI startup backed by top-tier investors, with adoption at companies like Instacart, Sofi, and Betterment. This role stands out for its focus on building and operating production AI systems in a fully in-person NYC HQ near Union Square.

175K–225K
awskubernetesdocker+2 more