SciforiumVerified Source

GPU Cluster Engineer Networking Sciforium

150K–180K
Onsite · San Francisco, California
Posted July 24, 2026
payroll

Overview

Senior Network Engineer leading GPU cluster networking at Sciforium. Own full stack from RDMA fabric to cloud connectivity. Design and operate high-performance networks for large-scale AI workloads. Differentiate by working directly with AMD engineers and scaling cutting-edge infrastructure.

What You'll Do15

  • 1Design and architect full network for new GPU clusters including compute backend fabric storage network and perimeter networks
  • 2Select and justify InfiniBand versus RoCE fabrics for training and inference workloads
  • 3Specify hardware components such as switches NICs and optics while producing port maps and cable schedules
  • 4Own logical network design including IP addressing VLAN segmentation and BGP/EVPN-VXLAN architecture
  • 5Configure and validate RoCE v2 lossless Ethernet with congestion control and QoS policies
  • 6Run subnet managers adaptive routing and SHARP aggregation for optimal performance
  • 7Validate new fabrics using benchmarking tools and verify GPUDirect RDMA paths
  • 8Debug fabric-level issues like congestion trees and link flaps through deep packet analysis
  • 9Maintain BGP OSPF ECMP and EVPN-VXLAN across multi-vendor network operating systems
  • 10Manage firewall VPNs NAT and segmentation between production research and management domains
  • 11Generate network configurations from source of truth using Ansible/Nornir/NAPALM with CI validation
  • 12Implement telemetry streaming gNMI sFlow and optics monitoring into observability platforms
  • 13Perform switch and NIC firmware upgrades with minimal workload disruption
  • 14Design dark fiber DWDM inter-site links and hybrid cloud connectivity solutions
  • 15Engineer WAN QoS and traffic paths for cross-cluster replication and checkpoint traffic

Requirements12

  • 17+ years designing and operating production data center networks including large-scale HPC AI clusters
  • 2Expert routing switching including BGP EVPN-VXLAN OSPF ECMP VRF across multiple vendors
  • 3Deep RDMA expertise in RoCE v2 tuning InfiniBand fabric management and SHARP implementation
  • 4Hands-on experience with 100-800G optics high-radix switching and Clos/rail-optimized topologies
  • 5Production network security experience with enterprise firewalls and segmentation design
  • 6Network automation proficiency using Python Ansible or Nornir/NAPALM with Git workflows
  • 7Knowledge of host-side RDMA stack including MOFED DOCA NIC tuning and GPUDirect RDMA
  • 8Cloud networking experience with VPC design and dedicated interconnects
  • 9Dark fiber DWDM procurement and operations background preferred
  • 10BlueField DPU or SmartNIC deployment experience beneficial
  • 11Experience supporting distributed training at 1000+ GPU scale or multi-cluster operations
  • 12Expertise in Kubernetes networking for GPU serving with CNI SR-IOV Multus

Salary Insight

$150 - $180k per year

Location

Typeonsite
LocationSan Francisco, California

Required Skills

InfiniBandRoCE v2BGPEVPN-VXLANOSPFECMPVLANVRFPythonAnsibleNornirNAPALMNetBox
Share:

Similar open positions

Explore active roles that match your skills and interests.

Crusoe

Crusoe

19d agoSan Francisco, Californiapayroll

Senior Staff Software Engineer SDN Architecture

Crusoe seeks a Senior Staff Software Engineer to drive the software-defined networking stack enabling AI-first cloud infrastructure. This role involves designing high-performance networking solutions for massive GPU clusters while collaborating across product hardware and infrastructure teams.

Competitive salary
pythonawsreact+2 more
Celestica

Celestica

13h agoDallas, Texaspayroll

Senior Lead Software Engineer, GPU Data Centers

You will architect and validate a full stack application for next-generation data centers with GPU/AI compute elements. Build orchestration software for the entire rack, integrated visualization tools, and diagnostics to optimize GPU utilization. Collaborate with cross-functional teams to ship production-ready code and mentor engineers. This role stands out through its focus on Cloud Native methods, Kubernetes deployments, and GenAI tool adoption for development efficiency.

Competitive salary
PythonGoKubernetes+7 more

Crusoe

6d agoSan Francisco, Californiapayroll

Senior Staff Network Architect at Crusoe

Crusoe seeks a Senior Staff Network Architect to define and evolve network architecture across its AI infrastructure stack from bare-metal data center switching fabrics to SDN control planes and cloud-layer overlays. This high-impact role reports to the Director of Networking and drives adoption of intent-based networking and declarative configuration management. The position offers the chance to shape the architecture of tomorrow's AI infrastructure at a unique hyperscale campus.

Competitive salary
BGPEVPN/VXLANOSPF+5 more
Crusoe

Crusoe

17h agoSan Jose, Californiapayroll

Senior Staff Network Architect Crusoe

Crusoe seeks a Senior Staff Network Architect to define and evolve network architecture across its infrastructure stack from bare-metal data center switching fabrics to SDN control planes and cloud-layer overlays. This high-impact role reports to the Director of Networking and drives the design of network solutions for a 1.2 GW hyperscale campus. The position focuses on scaling modern networking paradigms and mentoring senior engineers.

Competitive salary
BGPEVPN/VXLANOSPF+4 more

Crusoe

4d agoSan Francisco, Californiapayroll

Principal Systems Software Engineer Crusoe

Crusoe seeks a Principal Systems Software Engineer to lead the visionary development of next-generation AI infrastructure. This role involves bridging silicon and software to unify Bare-Metal-as-a-Service Intelligent IaaS and Elastic CaaS into a high-performance pool of intelligence. The ideal candidate thrives in a fast-paced environment and drives massive-scale training workloads to hardware limits.

260K–340K
Linux kernelKVMQEMU+6 more

Introl Solutions LLC

16h agoDenver, Coloradocontract

Data Center Technician - GPU Infrastructure Deployment

Introl seeks skilled Data Center Technicians to join our GPU infrastructure deployment teams. This contract role offers autonomy while contributing to large-scale compute projects across the United States. Ideal candidates thrive in demanding environments and deliver measurable results.

22K–35K
GPU infrastructure deploymentHardware diagnosticsFiber optic installation+7 more