Senior Site Reliability Engineer - Infra Ops at Circle
Overview
Join Circle as a Senior Site Reliability Engineer to design and operate secure scalable platforms for digital assets and AI workloads. You will lead on‑call incidents and drive reliability improvements while collaborating with cross‑functional teams. This role offers a unique chance to shape the reliability strategy for a globally‑focused financial technology platform.
What You'll Do11
- 1Design and operate Kubernetes platforms that deliver secure highly available services across hybrid and public clouds
- 2Create reusable infrastructure as code with Terraform to automate provisioning and enforce governance
- 3Develop backend services and automation in Go Python or JavaScript/TypeScript to boost developer productivity
- 4Partner with product and engineering to translate workload needs into resilient designs
- 5Own production reliability through on‑call leadership incident response and root‑cause analysis
- 6Establish SLIs SLOs error budgets and disaster recovery plans to meet customer expectations
- 7Embed security and compliance practices with the Security team to protect data and meet regulations
- 8Apply AI‑assisted tooling to enhance observability and reduce alert noise
- 9Mentor team members foster knowledge sharing and collaborative growth
- 10Drive continuous improvement of CI/CD pipelines and deployment strategies
- 11Contribute to a culture of high integrity and future forward thinking
Requirements11
- 15+ years of experience in Site Reliability Engineering DevOps or Infrastructure Engineering
- 2Deep hands‑on Kubernetes expertise including cluster design and troubleshooting at scale
- 3Strong Terraform background with reusable modules and state management experience
- 4Production software development skills in Go Python or JavaScript/TypeScript
- 5Proven track record improving reliability performance or cost efficiency of distributed systems
- 6Familiarity with cloud infrastructure networking IAM DNS load balancing and secure connectivity
- 7Experience defining and enforcing SLIs SLOs error budgets and incident management processes
- 8Knowledge of CI/CD GitOps or canary blue‑green deployment practices
- 9Security‑focused mindset and collaboration with Security teams in regulated environments
- 10Clear written and verbal communication with strong ownership and judgment
- 11Experience using AI‑assisted tooling is a plus
Salary Insight
Salary not disclosed in listing
Location
Required Skills
Similar open positions
Explore active roles that match your skills and interests.
Circle Internet Management Services LLC
VerifiedSite Reliability Engineer Circle Crypto Blockchain Infrastructure
Circle seeks a Staff Site Reliability Engineer to design and operate blockchain infrastructure at scale. The role involves owning reliability and performance of distributed systems across public clouds while collaborating with cross-functional teams. This position emphasizes high ownership and driving technical excellence.
C0035 LiveRamp, Inc.
VerifiedSenior SRE (Site Reliability Engineer) - LiveRamp
You will own the deployment and reliability of global products at LiveRamp, the data collaboration platform used by hundreds of innovative companies. You will set up production and internal environments, provide 24/7 first-line support, and drive resolutions with engineering teams. You will enhance CI/CD tooling, maintain Terraform scripts, and optimize system performance and cost. This role stands out for its global scope, working with teams across California, Paris, Nantong, Singapore, and Australia, and for championing SRE best practices.

Dice Talent & Staffing Solutions
VerifiedSite Reliability Engineer, SRE & Cloud Infrastructure
You will own the reliability and scalability of mission-critical systems for a cutting-edge government contractor in San Francisco. You will design, build, and run large-scale distributed systems with AWS, Kubernetes, and Terraform. You will work with a tight-knit team of engineers to ensure 99.99% uptime for services that matter. This role demands a hands-on engineer who thrives on incident response and automation. You will directly shape the infrastructure strategy from day one.
Cisco Systems
VerifiedStaff Site Reliability Engineer SRE Cisco Systems Hybrid
Lead technical roadmap for platform reliability scalability and operational excellence. Drive architecture evolution for cloud and air-gapped environments. Establish SLOs and resilience reviews. Lead major initiatives across Kubernetes databases and networking. Mentor engineers and collaborate with cross-functional teams to deliver secure highly reliable deployments.

Xoriant Corporation
VerifiedSenior/Staff SRE, AI/ML Platform Infrastructure
Own the reliability of a large-scale AI/ML platform serving millions of requests daily. Kubernetes and Docker are your primary tools. Join a team of 8 SREs supporting 50+ microservices on AWS and GCP. This role focuses on incident command, automation, and platform improvements.

Incorporan Inc
VerifiedSite Reliability Engineer, SRE & ML Systems
You will own the reliability, scalability, and performance of production systems at a global scale. You will work alongside platform and data engineering teams, integrating ML workloads with Python and TensorFlow into robust SRE practices. This role demands a blend of software engineering and systems thinking, operating in a hybrid environment with clear ownership of uptime and incident response.