Circle Internet Management Services LLCVerified Source
Remote

Senior Site Reliability Engineer - Infra Ops at Circle

Remote · San Francisco, California
Posted August 6, 2026
payroll

Overview

Join Circle as a Senior Site Reliability Engineer to design and operate secure scalable platforms for digital assets and AI workloads. You will lead on‑call incidents and drive reliability improvements while collaborating with cross‑functional teams. This role offers a unique chance to shape the reliability strategy for a globally‑focused financial technology platform.

What You'll Do11

  • 1Design and operate Kubernetes platforms that deliver secure highly available services across hybrid and public clouds
  • 2Create reusable infrastructure as code with Terraform to automate provisioning and enforce governance
  • 3Develop backend services and automation in Go Python or JavaScript/TypeScript to boost developer productivity
  • 4Partner with product and engineering to translate workload needs into resilient designs
  • 5Own production reliability through on‑call leadership incident response and root‑cause analysis
  • 6Establish SLIs SLOs error budgets and disaster recovery plans to meet customer expectations
  • 7Embed security and compliance practices with the Security team to protect data and meet regulations
  • 8Apply AI‑assisted tooling to enhance observability and reduce alert noise
  • 9Mentor team members foster knowledge sharing and collaborative growth
  • 10Drive continuous improvement of CI/CD pipelines and deployment strategies
  • 11Contribute to a culture of high integrity and future forward thinking

Requirements11

  • 15+ years of experience in Site Reliability Engineering DevOps or Infrastructure Engineering
  • 2Deep hands‑on Kubernetes expertise including cluster design and troubleshooting at scale
  • 3Strong Terraform background with reusable modules and state management experience
  • 4Production software development skills in Go Python or JavaScript/TypeScript
  • 5Proven track record improving reliability performance or cost efficiency of distributed systems
  • 6Familiarity with cloud infrastructure networking IAM DNS load balancing and secure connectivity
  • 7Experience defining and enforcing SLIs SLOs error budgets and incident management processes
  • 8Knowledge of CI/CD GitOps or canary blue‑green deployment practices
  • 9Security‑focused mindset and collaboration with Security teams in regulated environments
  • 10Clear written and verbal communication with strong ownership and judgment
  • 11Experience using AI‑assisted tooling is a plus

Salary Insight

Salary not disclosed in listing

Location

Typeremote
LocationSan Francisco, California
This is a remote position

Required Skills

KubernetesTerraformGoPythonJavaScript/TypeScriptCI/CDObservabilityCloud InfrastructureIAMDNSLoad Balancing
Share:

Similar open positions

Explore active roles that match your skills and interests.

Circle Internet Management Services LLC

Circle Internet Management Services LLC

6d agoRemotepayroll

Site Reliability Engineer Circle Crypto Blockchain Infrastructure

Circle seeks a Staff Site Reliability Engineer to design and operate blockchain infrastructure at scale. The role involves owning reliability and performance of distributed systems across public clouds while collaborating with cross-functional teams. This position emphasizes high ownership and driving technical excellence.

195K–258K
KubernetesTerraformPulumi+4 more
C0035 LiveRamp, Inc.

C0035 LiveRamp, Inc.

7d agoSan Francisco, Californiapayroll

Senior SRE (Site Reliability Engineer) - LiveRamp

You will own the deployment and reliability of global products at LiveRamp, the data collaboration platform used by hundreds of innovative companies. You will set up production and internal environments, provide 24/7 first-line support, and drive resolutions with engineering teams. You will enhance CI/CD tooling, maintain Terraform scripts, and optimize system performance and cost. This role stands out for its global scope, working with teams across California, Paris, Nantong, Singapore, and Australia, and for championing SRE best practices.

135K–159K
TerraformJenkinsCircleCI+10 more
Dice Talent & Staffing Solutions

Dice Talent & Staffing Solutions

10h agoSan Francisco, Californiapayroll

Site Reliability Engineer, SRE & Cloud Infrastructure

You will own the reliability and scalability of mission-critical systems for a cutting-edge government contractor in San Francisco. You will design, build, and run large-scale distributed systems with AWS, Kubernetes, and Terraform. You will work with a tight-knit team of engineers to ensure 99.99% uptime for services that matter. This role demands a hands-on engineer who thrives on incident response and automation. You will directly shape the infrastructure strategy from day one.

180K–265K
awskubernetesterraform+2 more
Cisco Systems

Cisco Systems

12d agoSan Francisco, Californiapayroll

Staff Site Reliability Engineer SRE Cisco Systems Hybrid

Lead technical roadmap for platform reliability scalability and operational excellence. Drive architecture evolution for cloud and air-gapped environments. Establish SLOs and resilience reviews. Lead major initiatives across Kubernetes databases and networking. Mentor engineers and collaborate with cross-functional teams to deliver secure highly reliable deployments.

187K–268K
KubernetesAWSGCP+8 more
Xoriant Corporation

Xoriant Corporation

1d agoSan Jose, Californiapayroll

Senior/Staff SRE, AI/ML Platform Infrastructure

Own the reliability of a large-scale AI/ML platform serving millions of requests daily. Kubernetes and Docker are your primary tools. Join a team of 8 SREs supporting 50+ microservices on AWS and GCP. This role focuses on incident command, automation, and platform improvements.

Competitive salary
Production on-callIncident commandBlameless postmortem+15 more
Incorporan Inc

Incorporan Inc

10h agoChicago, Illinoispayroll

Site Reliability Engineer, SRE & ML Systems

You will own the reliability, scalability, and performance of production systems at a global scale. You will work alongside platform and data engineering teams, integrating ML workloads with Python and TensorFlow into robust SRE practices. This role demands a blend of software engineering and systems thinking, operating in a hybrid environment with clear ownership of uptime and incident response.

104K–114K
pythontensorflowpytorch+2 more