GPU Cluster Engineer Networking Sciforium
Overview
Senior Network Engineer leading GPU cluster networking at Sciforium. Own full stack from RDMA fabric to cloud connectivity. Design and operate high-performance networks for large-scale AI workloads. Differentiate by working directly with AMD engineers and scaling cutting-edge infrastructure.
What You'll Do15
- 1Design and architect full network for new GPU clusters including compute backend fabric storage network and perimeter networks
- 2Select and justify InfiniBand versus RoCE fabrics for training and inference workloads
- 3Specify hardware components such as switches NICs and optics while producing port maps and cable schedules
- 4Own logical network design including IP addressing VLAN segmentation and BGP/EVPN-VXLAN architecture
- 5Configure and validate RoCE v2 lossless Ethernet with congestion control and QoS policies
- 6Run subnet managers adaptive routing and SHARP aggregation for optimal performance
- 7Validate new fabrics using benchmarking tools and verify GPUDirect RDMA paths
- 8Debug fabric-level issues like congestion trees and link flaps through deep packet analysis
- 9Maintain BGP OSPF ECMP and EVPN-VXLAN across multi-vendor network operating systems
- 10Manage firewall VPNs NAT and segmentation between production research and management domains
- 11Generate network configurations from source of truth using Ansible/Nornir/NAPALM with CI validation
- 12Implement telemetry streaming gNMI sFlow and optics monitoring into observability platforms
- 13Perform switch and NIC firmware upgrades with minimal workload disruption
- 14Design dark fiber DWDM inter-site links and hybrid cloud connectivity solutions
- 15Engineer WAN QoS and traffic paths for cross-cluster replication and checkpoint traffic
Requirements12
- 17+ years designing and operating production data center networks including large-scale HPC AI clusters
- 2Expert routing switching including BGP EVPN-VXLAN OSPF ECMP VRF across multiple vendors
- 3Deep RDMA expertise in RoCE v2 tuning InfiniBand fabric management and SHARP implementation
- 4Hands-on experience with 100-800G optics high-radix switching and Clos/rail-optimized topologies
- 5Production network security experience with enterprise firewalls and segmentation design
- 6Network automation proficiency using Python Ansible or Nornir/NAPALM with Git workflows
- 7Knowledge of host-side RDMA stack including MOFED DOCA NIC tuning and GPUDirect RDMA
- 8Cloud networking experience with VPC design and dedicated interconnects
- 9Dark fiber DWDM procurement and operations background preferred
- 10BlueField DPU or SmartNIC deployment experience beneficial
- 11Experience supporting distributed training at 1000+ GPU scale or multi-cluster operations
- 12Expertise in Kubernetes networking for GPU serving with CNI SR-IOV Multus
Salary Insight
$150 - $180k per year
Location
Required Skills
Similar open positions
Explore active roles that match your skills and interests.
Crusoe
VerifiedSenior Staff Software Engineer SDN Architecture
Crusoe seeks a Senior Staff Software Engineer to drive the software-defined networking stack enabling AI-first cloud infrastructure. This role involves designing high-performance networking solutions for massive GPU clusters while collaborating across product hardware and infrastructure teams.
Celestica
VerifiedSenior Lead Software Engineer, GPU Data Centers
You will architect and validate a full stack application for next-generation data centers with GPU/AI compute elements. Build orchestration software for the entire rack, integrated visualization tools, and diagnostics to optimize GPU utilization. Collaborate with cross-functional teams to ship production-ready code and mentor engineers. This role stands out through its focus on Cloud Native methods, Kubernetes deployments, and GenAI tool adoption for development efficiency.
Crusoe
VerifiedSenior Staff Network Architect at Crusoe
Crusoe seeks a Senior Staff Network Architect to define and evolve network architecture across its AI infrastructure stack from bare-metal data center switching fabrics to SDN control planes and cloud-layer overlays. This high-impact role reports to the Director of Networking and drives adoption of intent-based networking and declarative configuration management. The position offers the chance to shape the architecture of tomorrow's AI infrastructure at a unique hyperscale campus.
Crusoe
VerifiedSenior Staff Network Architect Crusoe
Crusoe seeks a Senior Staff Network Architect to define and evolve network architecture across its infrastructure stack from bare-metal data center switching fabrics to SDN control planes and cloud-layer overlays. This high-impact role reports to the Director of Networking and drives the design of network solutions for a 1.2 GW hyperscale campus. The position focuses on scaling modern networking paradigms and mentoring senior engineers.
Crusoe
VerifiedPrincipal Systems Software Engineer Crusoe
Crusoe seeks a Principal Systems Software Engineer to lead the visionary development of next-generation AI infrastructure. This role involves bridging silicon and software to unify Bare-Metal-as-a-Service Intelligent IaaS and Elastic CaaS into a high-performance pool of intelligence. The ideal candidate thrives in a fast-paced environment and drives massive-scale training workloads to hardware limits.
Introl Solutions LLC
VerifiedData Center Technician - GPU Infrastructure Deployment
Introl seeks skilled Data Center Technicians to join our GPU infrastructure deployment teams. This contract role offers autonomy while contributing to large-scale compute projects across the United States. Ideal candidates thrive in demanding environments and deliver measurable results.