Large Language Model Inference Engineer ByteDance
Overview
ByteDance seeks a graduate to engineer a large model inference system at Volcano Ark. You will build and scale high‑performance inference clusters while reducing costs through advanced techniques. This role drives innovation in AI infrastructure and impacts enterprise solutions across multiple ByteDance products.
What You'll Do11
- 1Design and implement the MaaS inference system for Volcano Ark
- 2Optimize large model inference performance cost and stability at scale
- 3Reduce inference costs via disaggregated multi‑role inference and heterogeneous inference
- 4Apply elastic computing and multi‑tenant co‑located inference strategies
- 5Collaborate with cross‑functional teams to ship features within 90 days
- 6Drive research on GPU performance analysis and low‑level system bottlenecks
- 7Maintain and evolve Kubernetes and Ray based resource orchestration frameworks
- 8Explore distributed network communication optimization using RDMA principles
- 9Contribute to Kubernetes and Ray frameworks for scalable workload management
- 10Support internal teams with custom ML compute and algorithm development
- 11Promote performance improvements through system‑level engineering practices
Requirements10
- 1Currently completing or recently graduated with a BS or MS in Computer Science or related field
- 2Strong grasp of algorithms design patterns data structures operating systems and computer architecture
- 3Proficiency in C++ or Python with solid coding standards
- 4Deep understanding of GPU hardware CUDA and performance analysis techniques
- 5Demonstrated interest in distributed systems large‑scale heterogeneous inference and low‑level performance study
- 6Experience with PD disaggregation and KV cache systems for multi‑node inference
- 7Expertise in distributed communication optimization RDMA operator implementation
- 8Familiarity with Kubernetes and Ray for resource orchestration and scheduling
- 9Graduate status or recent completion within the past year
- 10Ability to start onboarding by year‑end
Salary Insight
$128 - $256k per year
Location
Required Skills
Similar open positions
Explore active roles that match your skills and interests.
ByteDance
VerifiedLarge Language Model Training System Engineer Graduate Applied Machine Learning
ByteDance seeks a graduate to develop Volcano Ark training systems for large model post-training and reinforcement learning. This role involves designing elastic training solutions and optimizing distributed systems. The ideal candidate thrives in a fast-paced environment and contributes to cutting-edge AI infrastructure.
zoox
VerifiedMachine Learning Engineer Multi-Modality Foundation Model
Design and manage system capacity for perception models on autonomous vehicles. Lead initiatives to share perception stack components for efficient inference. Focus on delivering production-ready large-scale models for on-vehicle stacks. Seek experts experienced in compressing accelerating deploying complex models for power-constrained SoCs.
NUBYT, Inc.
VerifiedLLM Research Engineer IV at NUBYT Inc Mountain View
We seek a skilled LLM Research Engineer IV to design and fine-tune state-of-the-art Large Language Models. This role drives next-generation generative AI bridging cutting-edge NLP research and scalable production systems. The ideal candidate thrives in a collaborative environment and contributes directly to impactful outcomes.

Technogen, Inc.
VerifiedSenior LLM Inference & GPU Systems Engineer
Technogen, Inc. seeks a Senior On-Premise LLM Inference & GPU Systems Engineer to architect and optimize high-performance inference platforms in Charlotte, NC. You will own the deployment and tuning of large language models on on-premise GPU clusters, ensuring low-latency, high-throughput serving for enterprise workloads. Collaborate with data scientists and infrastructure teams to build scalable ML pipelines. This role offers direct impact on production AI systems within a Woman-Owned Small Business with 15+ years of IT services.
Intel
VerifiedAI Infrastructure Engineer Intel
Performance‑obsessed AI Infrastructure Engineer at Intel in San Jose, California. You will drive inference performance and redefine peak performance on Intel’s next‑generation GPU architectures.
Bain & Co.
VerifiedSenior AI/ML Engineer, LLMOps & RAG Systems
You build production inference, serving, and LLMOps infrastructure for Python-based ML systems, taking models from prototype to governed deployment. Bain's Private Equity Group Innovation team creates proprietary data and software products, and your work supports over 1,000 professionals across the investment lifecycle. You collaborate with Data Scientists, Data Engineers, and the Agent/AI squad to deliver reliable, observable ML systems. This hands-on role sets engineering standards and mentors mid-level engineers, with daily use of MLflow, Kubernetes, and Databricks.