Infrastructure Engineer at Cognition
Overview
Own agent execution infrastructure and develop the developer platform that powers Devin and Windsurf. Lead reliability scalability and craftsmanship for systems that enable AI agents to run at massive scale. This role demands deep systems engineering and ownership of compute networking and platform solutions.
What You'll Do7
- 1Design and operate sandboxed compute environments for agent task execution including VM orchestration container management and resource scheduling at scale
- 2Build and maintain developer platform with CI/CD pipelines deployment systems and internal tooling
- 3Define SLOs create monitoring and alerting systems lead incident response and improve observability
- 4Anticipate capacity needs and invest in infrastructure that scales with product growth
- 5Partner with product engineers to design systems that meet evolving requirements
- 6Maintain high reliability and security through sandboxing network isolation and multi-tenant best practices
- 7Collaborate across engineering to deliver seamless experiences for developers using AI tools
Requirements8
- 15+ years building large-scale distributed infrastructure with proven reliability track record
- 2Expertise in Kubernetes cloud platforms AWS GCP or Azure and IaC tools like Terraform
- 3Strong Python skills and experience owning complex codebases
- 4Deep knowledge of observability practices instrumentation and alert design
- 5Security and isolation mindset for sandboxed compute environments
- 6Relevant experience at frontier AI labs or applied AI companies
- 7BS MS or equivalent from top-tier university
- 8Candidate understands modern developer tooling and platform engineering
Salary Insight
$260 - $300k per year
Location
Required Skills
Similar open positions
Explore active roles that match your skills and interests.
Friendliai
VerifiedCloud Infrastructure Engineer, Kubernetes & AWS
FriendliAI seeks a Cloud Infrastructure Engineer to own the architecture and evolution of the GPU-accelerated AI inference cloud. You will design multi-cluster,Kubernetes fleets, extend the scheduler, and own the network path for latency-sensitive traffic. Work with the inference engine, platform, SRE, and security teams to turn serving demands into platform capabilities. This role offers a hands-on architecture position for an engineer ready to push large clusters further.
superblocks
VerifiedInfra Engineer - Execution Engine
Design and operate scalable production systems supporting multi-tenant cloud and on-premise deployments. Build a real-time distributed execution engine that powers AI applications and agents. Partner with product and customers to define roadmaps and deliver new builder experiences. Lead the creation of generational AI infrastructure beyond a typical nine-to-five role.
David Joseph & Company
VerifiedFull-Stack Engineer AI Agent Infrastructure
Lead ownership of full-stack product engineering for AI agent monitoring platform. Drive end-to-end feature development from data layer to UI while collaborating with customers to solve complex behavioral challenges.
extremenetworks
VerifiedDirector of SW Engineering - Exchange
Extreme is a global networking leader delivering cloud-driven solutions trusted by 50,000+ customers. We foster inclusion and drive scalable outcomes through innovative AI and agentic systems. Join us to shape the future of intelligent networking.
Hi Marley
VerifiedSr. Software Engineer II DevOps
Lead design and operation of cloud infrastructure on AWS to support core SaaS platform and agentic AI services. Build AI/ML infrastructure and monitoring for LLM-powered services. Establish IaC standards using Terraform. Implement observability beyond availability. Support big data pipelines, warehousing, and analytics. Drive disaster recovery and improve infrastructure parity. Lead architecture reviews and innovate on developer experience. This role differs by focusing on scaling autonomous AI agents in regulated insurance workflows.
Tetrix
VerifiedSenior Infra/DevOps Engineer, AWS & CI/CD
Own infrastructure for a AWS-native platform serving institutional investors in private markets. You will design scalable systems for high-volume document ingestion and AI-powered data pipelines at scale. Work with a small team of engineers, reporting to the CTO, and collaborate directly with product and clients. This role stands out by combining infrastructure leadership with client-facing work and strategic input on product direction.