Rutgers UniversityVerified Source

IT Manager, High-Performance Computing (HPC)

124K–190K
Partially · New York, New York
Posted August 12, 2026
payroll

Overview

Lead Rutgers University's High-Performance Computing (HPC) platforms and closely related infrastructure, serving thousands of researchers. You will manage a team overseeing clustered compute, GPU resources, high-performance storage, and schedulers. Own the day-to-day reliability and long-term evolution of research computing. Drive operational excellence, collaborate with researchers and campus IT partners, and ensure secure, cost-effective services. This role coordinates vendor engagements and sets the technology roadmap, making a direct impact on research at scale.

What You'll Do8

  • 1Lead the team responsible for Rutgers HPC platforms and related infrastructure, ensuring reliable operations and strategic growth.
  • 2Oversee day-to-day reliability and long-term evolution of clustered compute, GPU resources, high-performance storage, and schedulers like Slurm.
  • 3Drive a culture of operational excellence, collaborating with researchers and campus IT partners to align services with research needs.
  • 4Ensure secure, performant, and cost-effective HPC services, planning capacity and guiding technology roadmaps to meet university priorities and compliance standards.
  • 5Coordinate vendor engagements and hardware lifecycle management to build a future-ready research computing ecosystem.
  • 6Implement monitoring and utilization tools (Prometheus/Grafana, XDMoD) to optimize cluster performance and inform capacity planning.
  • 7Champion security best practices, including identity integration (LDAP/AD), access controls, and patch/vulnerability management for shared compute and storage services.
  • 8Establish and refine operational processes, including incident response, change management, and documentation, to improve team efficiency and service quality.

Requirements9

  • 17+ years of IT systems administration/engineering, with direct experience in clustered or large-scale environments, including hardware, virtualization, and production support.
  • 23+ years of team lead experience, with proven ability to mentor and manage technical staff.
  • 3Expertise with HPC technologies: Linux at scale, job schedulers (Slurm), workload accounting, and performance tuning.
  • 4Experience with high-performance research storage, including parallel filesystems like IBM Spectrum Scale (GPFS) and Lustre, with knowledge of quotas, snapshots, and disaster recovery.
  • 5Hands-on experience with InfiniBand fabrics, NVIDIA GPUs/accelerators (including CUDA and NCCL), and node provisioning tools such as Warewulf.
  • 6Proficiency with automation and configuration management using Ansible and scripting languages (Python, Bash) for operational efficiency.
  • 7Strong knowledge of security best practices for shared compute/storage services, including LDAP/Active Directory, access controls, and vulnerability management.
  • 8Excellent communication and collaboration skills to work effectively across central IT, researchers, and vendors.
  • 9Proven project management skills in planning, execution, risk management, and delivery, with experience in cost modeling and utilization reporting preferred.

Salary Insight

$124 - $190k per year

Location

Typepartially
LocationNew York, New York

Required Skills

LinuxSlurmIBM Spectrum ScaleInfiniBandGPUAnsibleScriptingProject Management
Share:

Similar open positions

Explore active roles that match your skills and interests.

ASK Consulting

19h agoSacramento, Californiapayroll

Senior DevOps Engineer, HPC & EDA Infrastructure

This senior role owns HPC and EDA platform operations, including SLURM cluster administration and datacenter migrations, at an enterprise scale. You will join the IT Datacenter (ITDC) team, coordinating with storage, EDA, and identity teams to deliver critical infrastructure changes. Your work ensures seamless service continuity for semiconductor and scientific computing workloads. This engagement offers direct impact on high-stakes migrations and production stability.

125K–133K
slurmslesansible+2 more
AT&T

AT&T

6d agoWashington, District of Columbiapayroll

HPC Systems Administrator AT&T SA2 Government

We seek a skilled HPC Systems Administrator to lead sustainment of Linux and Windows platforms at a Washington DC government contract. Responsibilities include managing clusters, supporting SRE teams, and optimizing performance. This role differs by focusing on large-scale HPC installations and unique security clearance requirements.

98K–168K
LinuxWindowsUNIX+11 more

Halogen Engineering Group, Inc

16d agoWashington, District of Columbiapayroll

Software Engineer HPC Linux

Halogen Engineering Group seeks a Software Engineer specializing in HPC and Linux systems. This role involves designing and maintaining robust permission management for critical resources while driving scalable software solutions. The ideal candidate will collaborate within a team to deliver high-performance applications across diverse domains.

218K–225K
pythonlinuxrbac+2 more

AHU Technologies Inc

23h agoWashington, District of Columbiapayroll

Software Integration Engineer HPC

Lead execution and maintenance of automated integration and system testing processes across HPC environments. Script automation system-level integration and report results to stakeholders. Ensure reliability through performance functional redundancy and failover testing.

246K–253K
Linux CLIBash scriptingPython scripting+10 more
WFMathpe

WFMathpe

15d agoRemotepayroll

Senior Project Manager Federal HPC Washington DC Remote

Lead a remote team to deliver complex HPC solutions while managing multi‑region revenue streams and high‑risk legal matters. This role drives strategic direction and financial oversight across extensive operations. It offers unique exposure to cutting‑edge edge‑to‑cloud technologies and fosters innovation through cross‑functional collaboration.

120K–275K
Project ManagementPMP CertificationHPC Solutions+5 more

Sciforium

19d agoSan Francisco, Californiapayroll

GPU Cluster Engineer Networking Sciforium

Senior Network Engineer leading GPU cluster networking at Sciforium. Own full stack from RDMA fabric to cloud connectivity. Design and operate high-performance networks for large-scale AI workloads. Differentiate by working directly with AMD engineers and scaling cutting-edge infrastructure.

150K–180K
InfiniBandRoCE v2BGP+10 more