Site Reliability Engineer, Kubernetes & AI Infrastructure
Overview
You will own the AWS-backed multi-tenant cloud and on-premise infrastructure that runs hundreds of thousands of AI applications at Superblocks. You will join a team of engineers from Uber, Stripe, and Datadog who built systems like Kafka, Kibana, and Debezium. We are a fast-growing AI startup backed by top-tier investors, with adoption at companies like Instacart, Sofi, and Betterment. This role stands out for its focus on building and operating production AI systems in a fully in-person NYC HQ near Union Square.
What You'll Do7
- 1Architect and operate scalable production systems supporting multi-tenant cloud and on-premise deployments.
- 2Design and develop a real-time distributed execution engine that powers all AI applications, workflows, and agents.
- 3Build, deploy, and optimize AI agent architecture, guardrails, and evals.
- 4Partner with product and customers to define the roadmap and bring new builder and AI experiences to life.
- 5Scale infrastructure to handle hundreds of thousands of AI applications with low latency and high reliability.
- 6Debug and resolve production issues across containers, VMs, caches, and task queues.
- 7Drive the adoption of Kubernetes and Docker for containerized deployments across cloud and on-prem environments.
Requirements6
- 13+ years managing cloud-based production apps with deep knowledge of containers, VMs, caches, task queues, networking, and OS.
- 2Designed and deployed infrastructure in production at scale with Docker, Kubernetes, ECS/EKS, or Firecracker.
- 3Strong product sense focused on great user experiences and strategic thinking to meet market and customer needs.
- 4Built and operated production AI systems and familiar with AI inference techniques.
- 5Optimized language runtimes and enable cross-language integration (e.g., Go, Python, C), including customizing or building WASM compilers and runtimes.
- 6Experience with machine learning algorithms, platforms, and frameworks like PyTorch and TensorFlow.
Salary Insight
$175 - $225k per year
Similar open positions
Explore active roles that match your skills and interests.
superblocks
VerifiedBackend Engineer, AI Agents & Governance
You will own the core control plane that manages hundreds of thousands of AI applications, with direct impact on RBAC, auth, and AI guardrails. Collaborating with a team from Uber, Stripe, and Datadog, you will build multi-tenant systems that run across cloud and on-prem. This role stands out for its focus on AI agent performance and safety in a fully in-person NYC team.

Revature
VerifiedSoftware Engineer, Enterprise AI Solutions
You will own end-to-end delivery of enterprise-scale AI solutions for Fortune 500 clients, starting with a 90-day ramp-up on a high-visibility project. You will join a cross-functional squad of engineers, data scientists, and product managers in a remote-first environment. This role stands apart through its focus on building engineers ahead of the AI shift, with direct mentorship from senior architects.

New York Technology Partners
VerifiedAI Data Engineer, Python & Kubernetes
Own the design and delivery of production-grade AI data pipelines at New York Technology Partners. You will build and scale data infrastructure on Kubernetes, tackling complex ETL/ELT workflows across cloud and on-premise systems. Work with a tight-knit team of data engineers and analysts, driving architecture decisions that power business intelligence. This role stands out for its focus on AI-driven data applications and modern data platforms.
Friendliai
VerifiedCloud Infrastructure Engineer, Kubernetes & AWS
FriendliAI seeks a Cloud Infrastructure Engineer to own the architecture and evolution of the GPU-accelerated AI inference cloud. You will design multi-cluster,Kubernetes fleets, extend the scheduler, and own the network path for latency-sensitive traffic. Work with the inference engine, platform, SRE, and security teams to turn serving demands into platform capabilities. This role offers a hands-on architecture position for an engineer ready to push large clusters further.

SRI Tech Solutions
VerifiedLead Site Reliability Engineer SRE
Own the design and scaling of high availability cloud infrastructure for a Generative AI platform. Drive reliability and operational excellence while leading technical initiatives. Shape architecture and mentor engineers across SRE practices.

Xoriant Corporation
VerifiedSenior/Staff SRE, AI/ML Platform Infrastructure
Own the reliability of a large-scale AI/ML platform serving millions of requests daily. Kubernetes and Docker are your primary tools. Join a team of 8 SREs supporting 50+ microservices on AWS and GCP. This role focuses on incident command, automation, and platform improvements.