Enterprise AI Infrastructure Engineer at GRAIL
Overview
Senior Staff Software Development Engineer leads design and scaling of enterprise AI platform using AWS and Kubernetes. This role drives technical excellence in cloud infrastructure and AI governance while mentoring teams and shaping enterprise AI strategy. Based in Sunnyvale with potential visits to Menlo Park and Durham NC.
What You'll Do11
- 1Lead end-to-end design deployment monitoring scalable governed AI platform using Amazon EKS and AWS native services
- 2Design agentic AI workflows autonomous agents multi-agent systems using LangChain LangGraph AutoGen Claude Agent SDK
- 3Architect secure integrations via Model Context Protocol connecting AI platform with internal systems vector databases third party SaaS applications
- 4Build identity authorization zero trust token flows with Okta Auth0 custom JWT authorizers
- 5Implement deterministic policy controls Cedar role based access approval gates human in the loop checks at API gateway
- 6Develop isolated containerized runtime environments Kubernetes pods on Amazon EKS for secure AI model execution
- 7Establish audit trails observability using AWS CloudTrail OpenTelemetry tracking cost latency tool calls
- 8Collaborate with Product Security Regulatory teams translating requirements into compliant AI infrastructure solutions
- 9Troubleshoot complex issues cloud infrastructure Kubernetes networking network isolation PrivateLink agentic workflows
- 10Contribute to technology roadmaps AI infrastructure strategy platform evolution mentoring engineers IaC best practices
- 11Partner with Quality Regulatory Privacy Security teams ensuring compliance with regulatory requirements
Requirements15
- 1Bachelor's degree in Computer Science Software Engineering AI Cloud Computing equivalent Master's or PhD preferred
- 28-12 years software development cloud infrastructure experience AWS EKS Kubernetes networking VPC PrivateLink IAM KMS GenAI services
- 3Proven agentic AI development autonomous agents LLM orchestration frameworks LangChain LangGraph AutoGen Claude Agent SDK
- 4Model Context Protocol implementation or robust API integrations for LLMs
- 5Identity access management OAuth JWT integration with enterprise IdPs Okta Auth0
- 6Advanced Python TypeScript Go programming infrastructure-as-code Terraform AWS CDK
- 7Vector databases RAG architectures row level access controls OpenSearch FAISS pgvector
- 8CI/CD pipelines MLOps Helm Kubernetes ecosystem containerization modern observability stacks
- 9Regulatory standards cybersecurity ISO 27001 NIST SOC 2 HIPAA medical device regulations IVDD IVDR FDA 21 CFR 800 series FDA 21 CFR Part 11
- 10Cloud infrastructure containerized environments agentic AI secure distributed system design
- 11Problem solving analytical skills addressing ambiguous high impact technical challenges
- 12Leadership influence driving alignment across engineering security regulatory business stakeholders
- 13Communication skills explaining complex LLM behaviors infrastructure architectures security boundaries
- 14Mentoring engineers cloud AI talent AI safety prompt injection defenses deterministic policy enforcement
- 15Strategic thinking balancing long term vision near term delivery adaptability intellectual curiosity emerging agentic AI frameworks MCP cloud trends
Salary Insight
Salary not disclosed in listing
Location
Required Skills
Similar open positions
Explore active roles that match your skills and interests.
Hi Marley
VerifiedSr. Software Engineer II DevOps
Lead design and operation of cloud infrastructure on AWS to support core SaaS platform and agentic AI services. Build AI/ML infrastructure and monitoring for LLM-powered services. Establish IaC standards using Terraform. Implement observability beyond availability. Support big data pipelines, warehousing, and analytics. Drive disaster recovery and improve infrastructure parity. Lead architecture reviews and innovate on developer experience. This role differs by focusing on scaling autonomous AI agents in regulated insurance workflows.
Ntt-Data-Aivista
VerifiedMember of Technical Staff Applied AI Engineering
Design and deploy agentic systems that empower enterprise AI products to operate reliably at scale. Own the end-to-end lifecycle from prototype to production while collaborating with cross‑functional teams to solve complex governance challenges. This builder role offers ownership of outcomes and drives impact in a fast‑moving environment.
Oracle
VerifiedLead Principal AI Software Engineer Oracle
Senior Staff-level technical leadership role responsible for defining building next-generation AI systems on Oracle Cloud Infrastructure. Sets architecture and engineering direction for production-grade agentic AI platforms autonomous workflows scalable inference infrastructure enterprise AI applications used in large-scale business-critical environments. Requires proven engineer who translates ambiguous product and platform goals into durable technical strategy leads multi-team execution without direct authority remains deeply hands-on in design code reviews operations and incident follow-up.
Capital_One
VerifiedDistinguished AI Engineer - Agentic AI Platform Remote Eligible
Capital One seeks a Distinguished AI Engineer to design and ship agentic AI platforms that transform banking. You will own workflow frameworks, SDKs, and guardrails while scaling responsible AI across tenant needs. This role drives platform architecture, developer experience, and trust initiatives. Unique opportunity to shape enterprise AI at scale.
grailbio
VerifiedForward Deployed Engineer AI GRAIL
GRAIL seeks a skilled Forward Deployed Engineer to design and deploy AI solutions across multiple domains. This role drives scalable software development while collaborating with cross-functional teams. You will own AI-enabled applications and contribute to enterprise initiatives.

TechniPros, LLC
VerifiedGen AI Engineer, Cloud & Enterprise AI
Own cloud-native AI development and deployment at enterprise scale. Build and ship AI applications on Azure AI Foundry, AWS Bedrock, or Google Vertex AI. Integrate AI services into production systems. Use RAG architectures to deliver accurate, context-aware responses. You will set the technical direction for AI integration and drive measurable business outcomes.