
Principal Observability Platform Engineer, GPU & AI Infrastructure
Overview
You will own the architectural roadmap for our observability platform, covering GPU clusters, AI workloads, and the underlying infrastructure. You will define and build the platform that provides deep visibility into Kubernetes, Prometheus, and Grafana stacks, and lead a team of engineers to raise the engineering bar. This role shapes the technical direction and scales infrastructure ahead of business needs, directly impacting the reliability of AI and ML pipelines.
What You'll Do7
- 1Design and implement a unified observability platform that handles 5+ million metrics per second across Kubernetes clusters.
- 2Build self-service telemetry pipelines using OpenTelemetry and Prometheus, enabling teams to ship custom metrics and traces in minutes.
- 3Lead a cross-functional group to migrate legacy monitoring to Grafana dashboards, reducing mean time to detection by 40%.
- 4Debug complex performance issues across GPU clusters and AI workloads, optimizing resource utilization and latency.
- 5Scale the logging and tracing infrastructure to support 10 TB of daily logs and 500 TB of traces.
- 6Drive the roadmap for incident response tooling, integrating PagerDuty and Slack alerts.
- 7Mentor engineers in SRE best practices, fostering ownership and on-call excellence.
Requirements7
- 18+ years in platform engineering or SRE roles, with 3+ years leading technical initiatives.
- 2Expertise in Kubernetes and container orchestration, including cluster networking and storage.
- 3Deep knowledge of Prometheus, Grafana, and OpenTelemetry for metrics, logs, and traces.
- 4Experience with AWS and GCP infrastructure, including EC2, S3, and managed Kubernetes (EKS/GKE).
- 5Proficiency in Go or Python for building telemetry components and tooling.
- 6Track record of scaling observability systems to handle millions of metrics and terabytes of data.
- 7Strong communication skills to influence technical decisions across teams.
Salary Insight
$104 - $114k per year
Location
Required Skills
Similar open positions
Explore active roles that match your skills and interests.

Arkhya Tech
VerifiedPlatform Architect, Observability & Grafana
You will own the design and delivery of observability platforms for enterprise customers, scaling across Grafana and related stacks. You will lead client requirement discussions independently, translating functional and security needs into robust solutions. Working with a team of engineers, you will ensure production-grade implementations with a customer-first mindset. This role demands hands-on troubleshooting and the ability to drive architecture decisions end-to-end.

SmallArc, Inc
VerifiedPlatform Architect Senior Observability Platform Architect
Lead design and delivery of observability platforms for large scale systems. Own technical strategy and drive cross functional collaboration. Shape product roadmap based on insights from monitoring and logging data. Differentiate by delivering scalable solutions that improve system reliability and performance.

Xoriant Corporation
VerifiedSenior/Staff SRE, AI/ML Platform Infrastructure
Own the reliability of a large-scale AI/ML platform serving millions of requests daily. Kubernetes and Docker are your primary tools. Join a team of 8 SREs supporting 50+ microservices on AWS and GCP. This role focuses on incident command, automation, and platform improvements.
Serko Ltd
VerifiedPrincipal Engineer AI Platform Operations
Serko Ltd seeks a Principal Engineer to architect the AI Platform & Operations. You will define the long-term technical roadmap for AI products, establish engineering benchmarks, and optimize GPU compute efficiency. This role drives platform stability and reliability while mentoring senior engineers and championing internal developer platforms.
Celestica
VerifiedSenior Lead Software Engineer, GPU Data Centers
You will architect and validate a full stack application for next-generation data centers with GPU/AI compute elements. Build orchestration software for the entire rack, integrated visualization tools, and diagnostics to optimize GPU utilization. Collaborate with cross-functional teams to ship production-ready code and mentor engineers. This role stands out through its focus on Cloud Native methods, Kubernetes deployments, and GenAI tool adoption for development efficiency.
Virtasant
VerifiedSenior Platform Engineer Cloud Infrastructure Go
We seek a senior engineer to design and operate cloud-native platforms using Go. This role involves building Kubernetes clusters managing networking workload isolation and multi-region topologies implementing service mesh features like mTLS and traffic management optimizing containerized workloads developing Go services and middleware writing unit and integration tests leading incident response defining SLOs and building alerting systems managing infrastructure as code and CI/CD pipelines planning cloud migrations and ensuring observability through metrics and tracing collaborating with product and security teams.