
Senior On-Premise LLM Inference & GPU Systems Engineer
Overview
You will build and maintain large-scale on-prem LLM infrastructure on NVIDIA H200 GPU clusters, powering enterprise private GenAI workloads. You will manage production inference, including self-hosting Llama and other open-source models, within an OpenShift AI deployment ecosystem. Your stack includes Kubernetes, Docker, and GPU orchestration, with a focus on reliability and performance. This role is pivotal for scaling inference in a hybrid cloud environment.
What You'll Do8
- 1Build Kubernetes clusters to run LLM inference workloads on NVIDIA H200 GPUs.
- 2Deploy and configure OpenShift AI for enterprise GenAI inference services.
- 3Self-host open-source LLM models such as Llama in production.
- 4Profile and tune GPU utilization and memory for inference performance.
- 5Develop monitoring and alerting for GPU health and inference latency.
- 6Implement security policies for model access and data isolation.
- 7Automate cluster scaling and model updates using Helm and CI/CD pipelines.
- 8Debug and resolve inference runtime issues across NVIDIA CUDA and TensorRT.
Requirements8
- 15+ years operating Kubernetes in production environments.
- 23+ years hands-on with NVIDIA GPU clusters, including CUDA and TensorRT.
- 3Experience deploying LLM inference servers like vLLM or Triton.
- 4Strong expertise in OpenShift or OKD deployment and management.
- 5Proficiency in Python and Bash scripting for automation.
- 6Familiarity with Networking (CNI, service mesh) and storage for inference.
- 73+ years using Docker and container orchestration.
- 8Understanding of model quantization and optimization for GPU.
Salary Insight
$65 - $75k per year
Location
Required Skills
Similar open positions
Explore active roles that match your skills and interests.

Technogen, Inc.
VerifiedSenior LLM Inference & GPU Systems Engineer
Technogen, Inc. seeks a Senior On-Premise LLM Inference & GPU Systems Engineer to architect and optimize high-performance inference platforms in Charlotte, NC. You will own the deployment and tuning of large language models on on-premise GPU clusters, ensuring low-latency, high-throughput serving for enterprise workloads. Collaborate with data scientists and infrastructure teams to build scalable ML pipelines. This role offers direct impact on production AI systems within a Woman-Owned Small Business with 15+ years of IT services.
Lavendo
VerifiedAI Field Engineer - Infrastructure Scaling
Own end-to-end AI infrastructure scaling for enterprise and AI-native clients. Lead discovery and production deployments from initial contact to live customer environments. Partner with VP Engineering and CTO-level leaders to drive technical wins. Work remotely with occasional travel to US hubs.
Intel
VerifiedAI Infrastructure Engineer Intel
Performance‑obsessed AI Infrastructure Engineer at Intel in San Jose, California. You will drive inference performance and redefine peak performance on Intel’s next‑generation GPU architectures.
Bain & Co.
VerifiedSenior AI/ML Engineer, LLMOps & RAG Systems
You build production inference, serving, and LLMOps infrastructure for Python-based ML systems, taking models from prototype to governed deployment. Bain's Private Equity Group Innovation team creates proprietary data and software products, and your work supports over 1,000 professionals across the investment lifecycle. You collaborate with Data Scientists, Data Engineers, and the Agent/AI squad to deliver reliable, observable ML systems. This hands-on role sets engineering standards and mentors mid-level engineers, with daily use of MLflow, Kubernetes, and Databricks.
RELX Inc. Company
VerifiedSenior Machine Learning Engineer III
Lead the implementation and scaling of AI systems for legal products. Partner with Data Scientists to turn validated models into reliable high-performance customer-facing systems. Own system architecture infrastructure and productionization of ML/LLM solutions. Based in Raleigh NC hybrid fully remote.

THE TILTED CIRCLE LLC
VerifiedSoftware Engineer, Edge AI Systems & LLM Deployment
4-10 years of software engineering experience comes alive as you own the full lifecycle of training, evaluating, and deploying LLM agents to edge hardware that interfaces with sensors and effectors. You will join a newly-formed team under The Tilted Circle LLC in Seattle, WA, working in a hybrid setup. Your stack includes Python, C++, and PyTorch, with deployment targets on NVIDIA Jetson and other hardware-constrained devices. This role stands apart by letting you shape the architecture from day one, collaborating with domain experts to solve real-time inference challenges at the tactical edge.