Technogen, Inc.
Technogen, Inc.Verified Source

Senior LLM Inference & GPU Systems Engineer

687K
Onsite · Charlotte, North Carolina
Posted August 2, 2026
payroll

Overview

Technogen, Inc. seeks a Senior On-Premise LLM Inference & GPU Systems Engineer to architect and optimize high-performance inference platforms in Charlotte, NC. You will own the deployment and tuning of large language models on on-premise GPU clusters, ensuring low-latency, high-throughput serving for enterprise workloads. Collaborate with data scientists and infrastructure teams to build scalable ML pipelines. This role offers direct impact on production AI systems within a Woman-Owned Small Business with 15+ years of IT services.

What You'll Do7

  • 1Design and deploy on-premise LLM inference servers using NVIDIA Triton or equivalent, targeting sub-100ms latency.
  • 2Tune GPU utilization and memory allocation to maximize throughput across A100/H100 clusters.
  • 3Build automated pipelines for model quantization, distillation, and pruning to reduce inference cost.
  • 4Implement monitoring and logging for inference endpoints, using Prometheus and Grafana.
  • 5Collaborate with platform teams to integrate inference APIs with Kubernetes-based orchestration.
  • 6Debug performance bottlenecks in CUDA kernels and optimize I/O for large model weights.
  • 7Document architecture and runbooks to enable 24/7 reliability for mission-critical AI services.

Requirements7

  • 15+ years in systems engineering with a focus on GPU computing and LLM inference.
  • 2Deep expertise in NVIDIA CUDA, cuDNN, and TensorRT for model optimization.
  • 3Proven experience with Docker and Kubernetes in production environments.
  • 4Strong scripting skills in Python and Bash for automation and tooling.
  • 5Hands-on with ONNX Runtime or vLLM for efficient serving.
  • 63+ years managing on-premise GPU clusters with SLURM or similar schedulers.
  • 7Bachelor's degree in Computer Science, Engineering, or related field.

Salary Insight

$687k per year

Location

Typeonsite
LocationCharlotte, North Carolina

Required Skills

nvidia tritonkubernetespythontensorrtdocker
Share:

Similar open positions

Explore active roles that match your skills and interests.

Lavendo

19d agoRemotepayroll

AI Field Engineer - Infrastructure Scaling

Own end-to-end AI infrastructure scaling for enterprise and AI-native clients. Lead discovery and production deployments from initial contact to live customer environments. Partner with VP Engineering and CTO-level leaders to drive technical wins. Work remotely with occasional travel to US hubs.

176K–224K
PythonKubernetesGPU infrastructure+5 more
RELX Inc. Company

RELX Inc. Company

14h agoRaleigh, North Carolinapayroll

Senior Machine Learning Engineer III

Lead the implementation and scaling of AI systems for legal products. Partner with Data Scientists to turn validated models into reliable high-performance customer-facing systems. Own system architecture infrastructure and productionization of ML/LLM solutions. Based in Raleigh NC hybrid fully remote.

118K–220K
PythonRustGo+6 more
Intel

Intel

14h agoSan Jose, Californiapayroll

AI Infrastructure Engineer Intel

Performance‑obsessed AI Infrastructure Engineer at Intel in San Jose, California. You will drive inference performance and redefine peak performance on Intel’s next‑generation GPU architectures.

170K–315K
C++PythonGPU Computing+8 more
THE TILTED CIRCLE LLC

THE TILTED CIRCLE LLC

10h agoSeattle, Washingtonpayroll

Software Engineer, Edge AI Systems & LLM Deployment

4-10 years of software engineering experience comes alive as you own the full lifecycle of training, evaluating, and deploying LLM agents to edge hardware that interfaces with sensors and effectors. You will join a newly-formed team under The Tilted Circle LLC in Seattle, WA, working in a hybrid setup. Your stack includes Python, C++, and PyTorch, with deployment targets on NVIDIA Jetson and other hardware-constrained devices. This role stands apart by letting you shape the architecture from day one, collaborating with domain experts to solve real-time inference challenges at the tactical edge.

Competitive salary
software engineeringAI systemsLLM agents+2 more

NUBYT, Inc.

7d agoSan Jose, Californiapayroll

LLM Research Engineer IV at NUBYT Inc Mountain View

We seek a skilled LLM Research Engineer IV to design and fine-tune state-of-the-art Large Language Models. This role drives next-generation generative AI bridging cutting-edge NLP research and scalable production systems. The ideal candidate thrives in a collaborative environment and contributes directly to impactful outcomes.

250K–260K
PyTorchTensorFlowJAX+5 more
Advent Global Solutions, Inc.

Advent Global Solutions, Inc.

12h agoDallas, Texascontract

Sr. AI Engineer, LLM & Generative AI

You own the design, build, and deployment of production-grade AI/ML solutions using Python, SQL, and modern AI frameworks. You join a team of engineers and data scientists, shipping LLM and Generative AI features that impact core business processes. You collaborate with product and infrastructure teams in a contract-to-hire role, 3 days onsite in Irving, TX. Your work directly improves model accuracy and system reliability from day one.

60K–65K
PythonSQLAPI+5 more