
Data Center Technician, GPU Infrastructure
Overview
You own uptime and hardware repair in a GPU-dense HPC environment operating at full scale. You work within a 24/7 crew of technicians supporting NVIDIA clusters and Cisco networking gear. Shift flexibility and rapid incident response make this a role for hands-on problem solvers.
What You'll Do8
- 1Diagnose and replace failed GPU boards, NVMe drives, and CPU modules within 24 hours of ticket creation.
- 2Execute server provisioning and decommissioning workflows under Linux and BMC management tools.
- 3Run power and thermal sweeps to maintain safe operating bounds on 550kW racks.
- 4Document all changes in the ServiceNow ticketing system with hourly check-ins.
- 5Coordinate with remote NOC and facilities teams to resolve network and power alarms.
- 6Lead shift handoffs, summarizing open incidents and required follow-ups.
- 7Validate firmware updates and BIOS configurations before releasing nodes to production.
- 8Run cable repairs and label audits across the MDF and IDF zones.
Requirements8
- 12+ years as a Data Center Technician or 4+ years in a similar hardware repair role.
- 2Hands-on experience with GPU platforms, including NVIDIA HGX and DGX systems.
- 3Working knowledge of Linux commands and BMC interfaces such as IPMI and Redfish.
- 4Ability to lift 50 pounds and work on feet for full shifts.
- 5Familiarity with fiber and copper cabling standards, including Cat6 and OM4.
- 6Comfortable with rotating shifts, including nights and weekends.
- 7A+ or Network+ certification is required.
- 8Associate's degree in IT, electronics, or a related field preferred.
Salary Insight
$73 - $94k per year
Similar open positions
Explore active roles that match your skills and interests.
Introl Solutions LLC
VerifiedData Center Technician - GPU Infrastructure Deployment
Introl seeks skilled Data Center Technicians to join our GPU infrastructure deployment teams. This contract role offers autonomy while contributing to large-scale compute projects across the United States. Ideal candidates thrive in demanding environments and deliver measurable results.
Insight Global
VerifiedDatacenter Technician, Network Infrastructure Deployment
You will own the physical deployment of next-generation network infrastructure across multiple data center locations, handling rack-and-stack, cabling, and hardware troubleshooting with minimal supervision. You will collaborate with Technical Program Managers, internal teams, and external vendors to ensure smooth project execution. This 9-month contract-to-hire role in Atlanta, GA requires day or night shifts and offers $30/hour plus full benefits from day one. You will thrive in a dynamic environment where your hands-on expertise directly drives capacity expansion.
Celestica
VerifiedSenior Lead Software Engineer, GPU Data Centers
You will architect and validate a full stack application for next-generation data centers with GPU/AI compute elements. Build orchestration software for the entire rack, integrated visualization tools, and diagnostics to optimize GPU utilization. Collaborate with cross-functional teams to ship production-ready code and mentor engineers. This role stands out through its focus on Cloud Native methods, Kubernetes deployments, and GenAI tool adoption for development efficiency.

Cloud Destinations LLC
VerifiedL2 Data Center Technician
You will own on-site hardware lifecycle for a data center in Santa Clara, CA, including server installs, break-fix, and cabling. You'll work within a 24/7 operations team, collaborating with engineering and remote hands. This role demands shift flexibility and precision. You will use Linux, TCP/IP, and RAID every day.
Sciforium
VerifiedGPU Cluster Engineer Networking Sciforium
Senior Network Engineer leading GPU cluster networking at Sciforium. Own full stack from RDMA fabric to cloud connectivity. Design and operate high-performance networks for large-scale AI workloads. Differentiate by working directly with AMD engineers and scaling cutting-edge infrastructure.
Amazon Data Services, Inc.
VerifiedData Center Technician 3 DCO | Amazon Hardware Repair
As a Data Center Technician 3 on the Data Center Operations (DCO) team, you will own server and network hardware troubleshooting across Amazon's data centers in San Francisco. You will diagnose failures, replace components, and drive high-impact incident responses to keep infrastructure running. Working with IT hardware and network protocols, you'll ensure uptime for critical services. This role stands out for direct hands-on work in a global-scale environment with 24/7 shift flexibility.