Staff Site Reliability Engineer, Kubernetes & AWS Cloud
Overview
Own scalable production infrastructure powering Skydio Cloud platform at scale. Lead reliability for Kubernetes and AWS based systems across multi region deployments. Collaborate with cloud, software, and security teams to ensure availability during critical operations. This role focuses on building and operating a resilient cloud platform with strong emphasis on observability and automation.
What You'll Do10
- 1Build, operate, and troubleshoot production Kubernetes/EKS clusters and upgrades
- 2Manage AWS infrastructure including VPCs, subnets, load balancers, IAM, databases
- 3Define and maintain infrastructure using Terraform and CI/CD pipelines
- 4Build monitoring, alerting, and observability for critical infrastructure
- 5Automate operational tasks using Python or Go in production environments
- 6Participate in on-call rotations and respond to incidents to minimize downtime
- 7Scale infrastructure across new regions and deployment environments
- 8Diagnose production issues across Kubernetes, AWS, Linux, networking, and databases
- 9Own deployment systems such as Argo CD, Spinnaker, GitHub Actions, GitLab CI/CD, or Jenkins
- 10Improve reliability and performance through proactive problem detection and resolution
Requirements9
- 18+ years of experience in Site Reliability Engineering or related roles
- 2Strong hands-on experience operating Kubernetes and production clusters
- 3Proficiency with AWS fundamentals (VPCs, subnets, networking, load balancers, IAM, EKS)
- 4Production experience with Terraform or similar IaC tooling
- 5Experience owning CI/CD and deployment systems (Argo CD, Spinnaker, GitHub Actions, GitLab CI/CD, Jenkins)
- 6Experience diagnosing production infrastructure and networking problems
- 7Experience solving meaningful scaling or reliability challenges
- 8Proficiency automating operational tasks with Python or Go
- 9Experience with multi-region infrastructure or on-premises deployments (bonus)
Salary Insight
Salary not disclosed in listing
Location
Required Skills
Similar open positions
Explore active roles that match your skills and interests.

VLink Inc
VerifiedSenior Site Reliability Engineer, Cloud & Observability
Own the reliability of distributed systems at scale across multiple locations. You will design, implement, and operate highly available services that power critical customer experiences. You will collaborate with cross functional teams to drive reliability, performance, and efficiency improvements. This role offers the chance to work on cutting edge infrastructure projects in a fast growing company.

System One
VerifiedSite Reliability Engineer, Production Operations
You own production reliability for critical internal and external applications across Pittsburgh, Cleveland, and Dallas sites. You drive incident response, performance tuning, and automation to maintain high availability. You collaborate with production support and engineering teams to reduce downtime and improve system resilience. This role focuses on proactive monitoring and reliability engineering, not just firefighting.

Dice Talent & Staffing Solutions
VerifiedSite Reliability Engineer, SRE & Cloud Infrastructure
You will own the reliability and scalability of mission-critical systems for a cutting-edge government contractor in San Francisco. You will design, build, and run large-scale distributed systems with AWS, Kubernetes, and Terraform. You will work with a tight-knit team of engineers to ensure 99.99% uptime for services that matter. This role demands a hands-on engineer who thrives on incident response and automation. You will directly shape the infrastructure strategy from day one.

SRI Tech Solutions
VerifiedLead Site Reliability Engineer SRE
Own the design and scaling of high availability cloud infrastructure for a Generative AI platform. Drive reliability and operational excellence while leading technical initiatives. Shape architecture and mentor engineers across SRE practices.

PRIMUS Global Services Inc.
VerifiedSite Reliability Engineer Hybrid
Own and scale high availability production systems across hybrid environments. Lead development of monitoring solutions and automation scripts. Ensure system resilience and performance optimization. Collaborate with cross functional teams to drive reliability initiatives. Differentiate by focusing on real world SRE practices.
superblocks
VerifiedSite Reliability Engineer, Kubernetes & AI Infrastructure
You will own the AWS-backed multi-tenant cloud and on-premise infrastructure that runs hundreds of thousands of AI applications at Superblocks. You will join a team of engineers from Uber, Stripe, and Datadog who built systems like Kafka, Kibana, and Debezium. We are a fast-growing AI startup backed by top-tier investors, with adoption at companies like Instacart, Sofi, and Betterment. This role stands out for its focus on building and operating production AI systems in a fully in-person NYC HQ near Union Square.