IT Manager, High-Performance Computing (HPC)
Overview
Lead Rutgers University's High-Performance Computing (HPC) platforms and closely related infrastructure, serving thousands of researchers. You will manage a team overseeing clustered compute, GPU resources, high-performance storage, and schedulers. Own the day-to-day reliability and long-term evolution of research computing. Drive operational excellence, collaborate with researchers and campus IT partners, and ensure secure, cost-effective services. This role coordinates vendor engagements and sets the technology roadmap, making a direct impact on research at scale.
What You'll Do8
- 1Lead the team responsible for Rutgers HPC platforms and related infrastructure, ensuring reliable operations and strategic growth.
- 2Oversee day-to-day reliability and long-term evolution of clustered compute, GPU resources, high-performance storage, and schedulers like Slurm.
- 3Drive a culture of operational excellence, collaborating with researchers and campus IT partners to align services with research needs.
- 4Ensure secure, performant, and cost-effective HPC services, planning capacity and guiding technology roadmaps to meet university priorities and compliance standards.
- 5Coordinate vendor engagements and hardware lifecycle management to build a future-ready research computing ecosystem.
- 6Implement monitoring and utilization tools (Prometheus/Grafana, XDMoD) to optimize cluster performance and inform capacity planning.
- 7Champion security best practices, including identity integration (LDAP/AD), access controls, and patch/vulnerability management for shared compute and storage services.
- 8Establish and refine operational processes, including incident response, change management, and documentation, to improve team efficiency and service quality.
Requirements9
- 17+ years of IT systems administration/engineering, with direct experience in clustered or large-scale environments, including hardware, virtualization, and production support.
- 23+ years of team lead experience, with proven ability to mentor and manage technical staff.
- 3Expertise with HPC technologies: Linux at scale, job schedulers (Slurm), workload accounting, and performance tuning.
- 4Experience with high-performance research storage, including parallel filesystems like IBM Spectrum Scale (GPFS) and Lustre, with knowledge of quotas, snapshots, and disaster recovery.
- 5Hands-on experience with InfiniBand fabrics, NVIDIA GPUs/accelerators (including CUDA and NCCL), and node provisioning tools such as Warewulf.
- 6Proficiency with automation and configuration management using Ansible and scripting languages (Python, Bash) for operational efficiency.
- 7Strong knowledge of security best practices for shared compute/storage services, including LDAP/Active Directory, access controls, and vulnerability management.
- 8Excellent communication and collaboration skills to work effectively across central IT, researchers, and vendors.
- 9Proven project management skills in planning, execution, risk management, and delivery, with experience in cost modeling and utilization reporting preferred.
Salary Insight
$124 - $190k per year
Location
Required Skills
Similar open positions
Explore active roles that match your skills and interests.
ASK Consulting
VerifiedSenior DevOps Engineer, HPC & EDA Infrastructure
This senior role owns HPC and EDA platform operations, including SLURM cluster administration and datacenter migrations, at an enterprise scale. You will join the IT Datacenter (ITDC) team, coordinating with storage, EDA, and identity teams to deliver critical infrastructure changes. Your work ensures seamless service continuity for semiconductor and scientific computing workloads. This engagement offers direct impact on high-stakes migrations and production stability.
AT&T
VerifiedHPC Systems Administrator AT&T SA2 Government
We seek a skilled HPC Systems Administrator to lead sustainment of Linux and Windows platforms at a Washington DC government contract. Responsibilities include managing clusters, supporting SRE teams, and optimizing performance. This role differs by focusing on large-scale HPC installations and unique security clearance requirements.
Halogen Engineering Group, Inc
VerifiedSoftware Engineer HPC Linux
Halogen Engineering Group seeks a Software Engineer specializing in HPC and Linux systems. This role involves designing and maintaining robust permission management for critical resources while driving scalable software solutions. The ideal candidate will collaborate within a team to deliver high-performance applications across diverse domains.
AHU Technologies Inc
VerifiedSoftware Integration Engineer HPC
Lead execution and maintenance of automated integration and system testing processes across HPC environments. Script automation system-level integration and report results to stakeholders. Ensure reliability through performance functional redundancy and failover testing.
WFMathpe
VerifiedSenior Project Manager Federal HPC Washington DC Remote
Lead a remote team to deliver complex HPC solutions while managing multi‑region revenue streams and high‑risk legal matters. This role drives strategic direction and financial oversight across extensive operations. It offers unique exposure to cutting‑edge edge‑to‑cloud technologies and fosters innovation through cross‑functional collaboration.
Sciforium
VerifiedGPU Cluster Engineer Networking Sciforium
Senior Network Engineer leading GPU cluster networking at Sciforium. Own full stack from RDMA fabric to cloud connectivity. Design and operate high-performance networks for large-scale AI workloads. Differentiate by working directly with AMD engineers and scaling cutting-edge infrastructure.