
Senior LLM Inference & GPU Systems Engineer
Overview
Technogen, Inc. seeks a Senior On-Premise LLM Inference & GPU Systems Engineer to architect and optimize high-performance inference platforms in Charlotte, NC. You will own the deployment and tuning of large language models on on-premise GPU clusters, ensuring low-latency, high-throughput serving for enterprise workloads. Collaborate with data scientists and infrastructure teams to build scalable ML pipelines. This role offers direct impact on production AI systems within a Woman-Owned Small Business with 15+ years of IT services.
What You'll Do7
- 1Design and deploy on-premise LLM inference servers using NVIDIA Triton or equivalent, targeting sub-100ms latency.
- 2Tune GPU utilization and memory allocation to maximize throughput across A100/H100 clusters.
- 3Build automated pipelines for model quantization, distillation, and pruning to reduce inference cost.
- 4Implement monitoring and logging for inference endpoints, using Prometheus and Grafana.
- 5Collaborate with platform teams to integrate inference APIs with Kubernetes-based orchestration.
- 6Debug performance bottlenecks in CUDA kernels and optimize I/O for large model weights.
- 7Document architecture and runbooks to enable 24/7 reliability for mission-critical AI services.
Requirements7
- 15+ years in systems engineering with a focus on GPU computing and LLM inference.
- 2Deep expertise in NVIDIA CUDA, cuDNN, and TensorRT for model optimization.
- 3Proven experience with Docker and Kubernetes in production environments.
- 4Strong scripting skills in Python and Bash for automation and tooling.
- 5Hands-on with ONNX Runtime or vLLM for efficient serving.
- 63+ years managing on-premise GPU clusters with SLURM or similar schedulers.
- 7Bachelor's degree in Computer Science, Engineering, or related field.
Salary Insight
$687k per year
Location
Required Skills
Similar open positions
Explore active roles that match your skills and interests.
Lavendo
VerifiedAI Field Engineer - Infrastructure Scaling
Own end-to-end AI infrastructure scaling for enterprise and AI-native clients. Lead discovery and production deployments from initial contact to live customer environments. Partner with VP Engineering and CTO-level leaders to drive technical wins. Work remotely with occasional travel to US hubs.
RELX Inc. Company
VerifiedSenior Machine Learning Engineer III
Lead the implementation and scaling of AI systems for legal products. Partner with Data Scientists to turn validated models into reliable high-performance customer-facing systems. Own system architecture infrastructure and productionization of ML/LLM solutions. Based in Raleigh NC hybrid fully remote.
Intel
VerifiedAI Infrastructure Engineer Intel
Performance‑obsessed AI Infrastructure Engineer at Intel in San Jose, California. You will drive inference performance and redefine peak performance on Intel’s next‑generation GPU architectures.

THE TILTED CIRCLE LLC
VerifiedSoftware Engineer, Edge AI Systems & LLM Deployment
4-10 years of software engineering experience comes alive as you own the full lifecycle of training, evaluating, and deploying LLM agents to edge hardware that interfaces with sensors and effectors. You will join a newly-formed team under The Tilted Circle LLC in Seattle, WA, working in a hybrid setup. Your stack includes Python, C++, and PyTorch, with deployment targets on NVIDIA Jetson and other hardware-constrained devices. This role stands apart by letting you shape the architecture from day one, collaborating with domain experts to solve real-time inference challenges at the tactical edge.
NUBYT, Inc.
VerifiedLLM Research Engineer IV at NUBYT Inc Mountain View
We seek a skilled LLM Research Engineer IV to design and fine-tune state-of-the-art Large Language Models. This role drives next-generation generative AI bridging cutting-edge NLP research and scalable production systems. The ideal candidate thrives in a collaborative environment and contributes directly to impactful outcomes.

Advent Global Solutions, Inc.
VerifiedSr. AI Engineer, LLM & Generative AI
You own the design, build, and deployment of production-grade AI/ML solutions using Python, SQL, and modern AI frameworks. You join a team of engineers and data scientists, shipping LLM and Generative AI features that impact core business processes. You collaborate with product and infrastructure teams in a contract-to-hire role, 3 days onsite in Irving, TX. Your work directly improves model accuracy and system reliability from day one.