Recruitment.ai
Recruitment.aiVerified Source

Principal Observability Platform Engineer, GPU & AI Infrastructure

104K–114K
Onsite · San Francisco, California
Posted August 13, 2026
payroll

Overview

You will own the architectural roadmap for our observability platform, covering GPU clusters, AI workloads, and the underlying infrastructure. You will define and build the platform that provides deep visibility into Kubernetes, Prometheus, and Grafana stacks, and lead a team of engineers to raise the engineering bar. This role shapes the technical direction and scales infrastructure ahead of business needs, directly impacting the reliability of AI and ML pipelines.

What You'll Do7

  • 1Design and implement a unified observability platform that handles 5+ million metrics per second across Kubernetes clusters.
  • 2Build self-service telemetry pipelines using OpenTelemetry and Prometheus, enabling teams to ship custom metrics and traces in minutes.
  • 3Lead a cross-functional group to migrate legacy monitoring to Grafana dashboards, reducing mean time to detection by 40%.
  • 4Debug complex performance issues across GPU clusters and AI workloads, optimizing resource utilization and latency.
  • 5Scale the logging and tracing infrastructure to support 10 TB of daily logs and 500 TB of traces.
  • 6Drive the roadmap for incident response tooling, integrating PagerDuty and Slack alerts.
  • 7Mentor engineers in SRE best practices, fostering ownership and on-call excellence.

Requirements7

  • 18+ years in platform engineering or SRE roles, with 3+ years leading technical initiatives.
  • 2Expertise in Kubernetes and container orchestration, including cluster networking and storage.
  • 3Deep knowledge of Prometheus, Grafana, and OpenTelemetry for metrics, logs, and traces.
  • 4Experience with AWS and GCP infrastructure, including EC2, S3, and managed Kubernetes (EKS/GKE).
  • 5Proficiency in Go or Python for building telemetry components and tooling.
  • 6Track record of scaling observability systems to handle millions of metrics and terabytes of data.
  • 7Strong communication skills to influence technical decisions across teams.

Salary Insight

$104 - $114k per year

Location

Typeonsite
LocationSan Francisco, California

Required Skills

kubernetesprometheusgrafanaopentelemetryaws
Share:

Similar open positions

Explore active roles that match your skills and interests.

Arkhya Tech

Arkhya Tech

1d agoNew York, New Yorkpayroll

Platform Architect, Observability & Grafana

You will own the design and delivery of observability platforms for enterprise customers, scaling across Grafana and related stacks. You will lead client requirement discussions independently, translating functional and security needs into robust solutions. Working with a team of engineers, you will ensure production-grade implementations with a customer-first mindset. This role demands hands-on troubleshooting and the ability to drive architecture decisions end-to-end.

Competitive salary
GrafanaObservability
SmallArc, Inc

SmallArc, Inc

1d agoNew York, New Yorkpayroll

Platform Architect Senior Observability Platform Architect

Lead design and delivery of observability platforms for large scale systems. Own technical strategy and drive cross functional collaboration. Shape product roadmap based on insights from monitoring and logging data. Differentiate by delivering scalable solutions that improve system reliability and performance.

120K–130K
Observability (Grafana)Platform ArchitectureTroubleshooting+2 more
Xoriant Corporation

Xoriant Corporation

1d agoSan Jose, Californiapayroll

Senior/Staff SRE, AI/ML Platform Infrastructure

Own the reliability of a large-scale AI/ML platform serving millions of requests daily. Kubernetes and Docker are your primary tools. Join a team of 8 SREs supporting 50+ microservices on AWS and GCP. This role focuses on incident command, automation, and platform improvements.

Competitive salary
Production on-callIncident commandBlameless postmortem+15 more
Serko Ltd

Serko Ltd

1d agoRemotepayroll

Principal Engineer AI Platform Operations

Serko Ltd seeks a Principal Engineer to architect the AI Platform & Operations. You will define the long-term technical roadmap for AI products, establish engineering benchmarks, and optimize GPU compute efficiency. This role drives platform stability and reliability while mentoring senior engineers and championing internal developer platforms.

168K–230K
pythonsparkairflow+2 more
Celestica

Celestica

1d agoDallas, Texaspayroll

Senior Lead Software Engineer, GPU Data Centers

You will architect and validate a full stack application for next-generation data centers with GPU/AI compute elements. Build orchestration software for the entire rack, integrated visualization tools, and diagnostics to optimize GPU utilization. Collaborate with cross-functional teams to ship production-ready code and mentor engineers. This role stands out through its focus on Cloud Native methods, Kubernetes deployments, and GenAI tool adoption for development efficiency.

Competitive salary
PythonGoKubernetes+7 more

Virtasant

1d agoRemotepayroll

Senior Platform Engineer Cloud Infrastructure Go

We seek a senior engineer to design and operate cloud-native platforms using Go. This role involves building Kubernetes clusters managing networking workload isolation and multi-region topologies implementing service mesh features like mTLS and traffic management optimizing containerized workloads developing Go services and middleware writing unit and integration tests leading incident response defining SLOs and building alerting systems managing infrastructure as code and CI/CD pipelines planning cloud migrations and ensuring observability through metrics and tracing collaborating with product and security teams.

Competitive salary
GoKubernetesTerraform+10 more